Budget: 4000 UAH Deadline: 5 days
Good day!
I will collect a database in Python, but not by blindly parsing websites, rather primarily through open academic APIs — OpenAlex, Crossref, ORCID, and OAI-PMH of Ukrainian journals on OJS. The data is provided officially and in batches: author, institution, position, source link, and in the metadata of articles in Crossref, the author's email for correspondence is included. I select department websites and institutional repositories from the top to cover those who publish little abroad.
The output will be one file in Excel and CSV: full name, email, institution, position, source link, collection date. I clean duplicates in two passes — by email and by normalized full name along with the institution, as people change affiliations and simple string comparison does not work here. I run emails through syntax and MX record checks, and I present the status in a separate column so you can immediately see what is valid from the database.
The question is: what volume benchmark are you aiming for — roughly 5,000 contacts or closer to 30-40,000? And should I leave rows without found emails as placeholders, or do you only need contacts with emails? This will determine how deep I need to go in the collection.
The price is 4,000 UAH, as in the budget, and the timeframe is 5 days. I will show an intermediate cut of a few hundred rows on the second day so you can verify the column structure before I proceed with the entire dataset.
Petro Pankov, BotCraft Group