Budget: 25000 USD Deadline: 14 days
We can take on such a system. The benchmark for the first working stage is from 45,000 UAH and 10-14 days. This is not just a parser; the key aspects are the quality of matches, deduplication, control of false profiles, and a proper structure of results in JSON or CSV =)
Based on experience, we have created data enrichment systems, searches through open sources, automation of collection, internal CRMs, and analytical pipelines. For this task, I would use Python, Playwright, or Scrapy, a separate search module through search engines, a processing queue, cache, verification rules, and scoring matches by name, company, address, state, website, and phone number.
I see the approach as follows:
> we take a small sample of your records and create a search prototype
> separately search for personal profiles, business pages, company websites, and available contacts
> each found match receives a trust score to avoid mixing people with the same names
> we deliver the result in a structure with sources, trust level, verification date, and reason for the match
Look, there’s a nuance here - LinkedIn and Facebook have restrictions on automated collection, so I wouldn’t build a solution on a fragile account entry. It’s better to combine search results, open pages, company websites, business directories, and attribute verification. This way, the system will be more stable, rather than like a house of cards in the wind.
Please clarify:
> what is the volume of the database at the first stage - 1,000, 50,000, or more records
> what is the acceptable error margin and what is more important - more found contacts or fewer false matches
Relevant examples from Ingello:
> https://business.ingello.com/vorfahr - automation and complex data processing for business processes
> https://business.ingello.com/fractal - agency approach and automation of complex workflows
> https://business.ingello.com/forma-crm - corporate system with data, roles, and structured logic
Main page for FLH - https://systems-fl.ingello.com/ua
After sampling 100-300 records, it will be possible to more accurately estimate the total budget for the entire dataset. Usually, the pilot shows the real quality of sources and prevents spending the budget on beautiful but blind automation.