• Projects 4
  • Rating 5.0
  • Rating 1 189

Budget: 4000 UAH Deadline: 5 days

Good day!

I will collect a database in Python, but not by blindly parsing websites, rather primarily through open academic APIs — OpenAlex, Crossref, ORCID, and OAI-PMH of Ukrainian journals on OJS. The data is provided officially and in batches: author, institution, position, source link, and in the metadata of articles in Crossref, the author's email for correspondence is included. I select department websites and institutional repositories from the top to cover those who publish little abroad.

The output will be one file in Excel and CSV: full name, email, institution, position, source link, collection date. I clean duplicates in two passes — by email and by normalized full name along with the institution, as people change affiliations and simple string comparison does not work here. I run emails through syntax and MX record checks, and I present the status in a separate column so you can immediately see what is valid from the database.

The question is: what volume benchmark are you aiming for — roughly 5,000 contacts or closer to 30-40,000? And should I leave rows without found emails as placeholders, or do you only need contacts with emails? This will determine how deep I need to go in the collection.

The price is 4,000 UAH, as in the budget, and the timeframe is 5 days. I will show an intermediate cut of a few hundred rows on the second day so you can verify the column structure before I proceed with the entire dataset.

  • Projects 3
  • Rating 3.5
  • Rating 572

Budget: 4000 UAH Deadline: 4 days

Hello, I'll start right away: on the first day, I will provide the structure in Excel/CSV and the first sample with links to sources. I will gather publicly available contacts through OpenAlex, ORCID, and official institution websites, normalize the fields, remove duplicates, and check the format and domain of the email. The result will be a clean table with full names, positions, institutions, and sources. The deadline is 4 days. Are there any priority institutions or sectors?

  • Projects 5
  • Rating 4.9
  • Rating 756

Budget: 4000 UAH Deadline: 7 days

Hello, I recently created a similar parser for academic sources: collecting full names, emails, workplaces, and positions from department pages and faculty profiles, with a reference to the source for each contact. Deduplication and email validity checking through API were also part of that project.

If the database is needed for mailings, it's worth running the emails through SMTP validation right away and filtering out non-working ones, which increases deliverability by about 20-30%.

How many contacts are approximately needed at the start, and is this a one-time collection or are regular updates needed? It's also important to know which sources are a priority - ORCID and OpenAlex provide cleaner data than ResearchGate.

I suggest we get in touch; I will also outline a parsing scheme based on the sources and the structure of the final CSV according to your fields.

  • Projects 7
  • Rating 5.0
  • Rating 404

Budget: 4000 UAH Deadline: 5 days

Hello! I have experience in collecting large databases from open sources and the internet.

I will compile a clean database without duplicates in Excel/CSV format with all fields and links to sources.

Feel free to contact me, I would be happy to collaborate.

  • Projects -
  • Rating -
  • Rating 355

Budget: 4000 UAH Deadline: 4 days

Good day, Zoryana. ORCID, OpenAlex, and Crossref will provide a clean structure of full names, affiliations, and positions with almost no manual work, but there is no personal email as a separate field there. The real source of email is the department page or the teacher's profile on the university website and the "for correspondence" block on the first page of the PDF of the article itself. Ukrainian journals on OJS additionally open metadata through OAI-PMH, which is faster than manually browsing websites and covers publications that are not in international databases.

Google Scholar and ResearchGate prohibit automated collection by their terms of use, so I take these two sources manually and in small batches, rather than as a main flow. I remove duplicates by email and separately by normalized full name along with the institution, as people's affiliations change and matching only by name is unreliable. I check the email for format and MX record of the domain, not for actual deliverability, as these are different things and the latter cannot be verified without sending an email.

For 4000 UAH and 4 days, I will collect the first batch of 2000-3000 unique verified contacts with a source link in each line, and on the second day, I will show a few hundred lines for structure verification. If after this you need to go to thirty-forty thousand, I will continue in separate batches at a price agreed upon for a thousand new contacts. What volume benchmark are you keeping, a few thousand or immediately such a scale?

  • Projects -
  • Rating -
  • Rating 278

Budget: 4000 UAH Deadline: 7 days

You need to compile an up-to-date database of contacts for Ukrainian scientists, educators, and graduate students from open academic sources in Excel or CSV format. I will set up data collection from university websites, department pages, ORCID, OpenAlex, Google Scholar, Crossref, and institutional repositories using Python and parsing. Each entry will include full name, email, workplace, position, and a link to the source. I will also add deduplication and check the validity of addresses where possible. I will build the automation to comply with service rules, and I will make the table user-friendly. I am ready to take on the task and gather a complete database from open sources for you. Please write to me to agree on the details and the required sample size.

  • Projects 56
  • Rating 5.0
  • Rating 1 754

Budget: 3999 UAH Deadline: 2 days

I have almost ready database of teachers who offer tutoring services. Screenshot of the database https://ibb.co/xSz30ZBS Number of contacts 30+ thousand But there are no fields for position, place of work, or exact full name.

  • Projects -
  • Rating -
  • Rating 304

Budget: 3900 UAH Deadline: 3 days

Hello, I am ready to set up the collection of contacts of Ukrainian scientists from academic resources and create a structured database in CSV format. Developing an algorithm for processing the pages of departments and specialized portals is a clear task. What approximate volume of the database do you plan to obtain in the first stage? I would be happy to receive your feedback, I will answer all your questions, and we can better orient ourselves regarding deadlines and other project details. I will gladly help.

  • Projects -
  • Rating -
  • Rating 273

Budget: 4000 UAH Deadline: 7 days

Good day! I have experience in automation and parsing open sources (Python, working with APIs, Google Sheets). I propose a pipeline: mass collection of names and affiliations through the OpenAlex and Crossref APIs for Ukrainian scientific institutions, supplementing emails through parsing PDF articles (corresponding author) and department/faculty pages on university websites, deduplication by full name. The result will be Excel/CSV: full name, email, workplace, position, source.

Since the volume of the "large database" in the terms of reference is not specified, I suggest starting with the collection of 1000-1500 unique valid contacts within 5 days and within the budget — if necessary, I will continue the collection as a separate stage. I am ready to discuss the details.

  • Projects 78
  • Rating 4.8
  • Rating 3 000

Budget: 4000 UAH Deadline: 2 days

Good day! I am ready to compile such a large and relevant contact database in just a couple of days! All contacts will be current and relevant!!! Feel free to reach out!!!

  • Projects 13
  • Rating 4.5
  • Rating 622

Budget: 4000 UAH Deadline: 7 days

Hello! I have experience in developing custom parsers and working with academic APIs (OpenAlex, Crossref, ORCID) and university websites. I am ready to compile a clean database of Ukrainian scientists with deduplication, links to sources, and export to Excel/CSV. How many records do you plan to collect?

  • Projects 3
  • Rating 5.0
  • Rating 778

Budget: 4000 UAH Deadline: 1 day

Hello! I am ready to implement this project. I have experience in writing custom scripts in Python for collecting, parsing, and cleaning data from open academic sources (OpenAlex, APIs of scientific platforms, university websites).

I will do everything through code (Python + BeautifulSoup/Pandas/API), set up deduplication and basic domain verification (MX records), so that you receive a clean table in CSV/Excel with the necessary fields and sources.

I am ready to discuss the details (what volume of the database is needed and if there are priority institutions) and agree on the deadlines and cost. Feel free to message me privately!

  • Projects 43
  • Rating 5.0
  • Rating 3 182

Budget: 4000 UAH Deadline: 10 days

Good day, I see you have reposted the project. I am resubmitting my services; I have already provided a sample of the data.

  • Projects -
  • Rating -
  • Rating 266

Budget: 4000 UAH Deadline: 1 day

Hello! I'm ready to complete the task, but it's a pity that I won't be chosen! Without reviews, everyone thinks I'm a bad specialist! Although I definitely do my job better than 99%. I can finish it in 1 day if the deadlines are of interest to you!

  • Projects 3
  • Rating 5.0
  • Rating 543

Budget: 4000 UAH Deadline: 2 days

Good day, Zoryana!

I will collect a database from official APIs — ORCID, OpenAlex, Crossref — plus the websites of departments of Ukrainian universities. The advantage of these sources is that they provide position and institution in a structured way, so the columns "position" and "place of work" will be filled, rather than empty, as can happen with regular page parsing.

I suggest starting with a pilot: 2000–3000 unique contacts with deduplication and verification of email domains, 4 days, 4000 UAH. You will see the quality and structure — then we can scale up, 1500 UAH for each subsequent thousand.

Honestly about two points. Google Scholar and ResearchGate prohibit automated collection, so I take them selectively by hand or not at all. And full verification of the "liveness" of emails is only possible through a test mailing — I check the correctness of the domain and format.

One question: what do you plan to use the database for? This will determine which fields are more important.

  • Projects 45
  • Rating 4.9
  • Rating 18 121

Budget: 6000 UAH Deadline: 5 days

Collecting a large database of scientists' emails is a task where the key is not just to find addresses, but to filter out duplicates and outdated contacts. I have experience in web scraping with Selenium and processing large volumes of data (Allegro API, auto auctions). Here’s how I see the solution:
1. Deduplication by full name and email, checking relevance through syntax and SMTP verification.
2. Export to CSV/Excel with links to sources.
I have worked on something similar — a Selenium scraper for an auto auction (https://freelancehunt.com/showcase/work/selenium-veb-skraper-dlya-izvlecheniya-dannyih/1874474.html). What is the volume of the database (thousands/tens of thousands)? Are there specific sources that need to be used? Is email validity verification required?

Price: 6000 UAH
Deadline: 5 days

  • Projects -
  • Rating -
  • Rating 342

Budget: 4000 UAH Deadline: 1 day

Hello.

I have fully reviewed your requirements and am ready to start as soon as we finish our discussion.

I will implement your requirements using a Python script.

  • Projects -
  • Rating -
  • Rating 367

Budget: 5000 UAH Deadline: 7 days

Stattiinua, welcome! I see that you need a database of emails of scientists and graduate students from open sources, with checks for duplicates and validity of addresses, exported to Excel/CSV with a reference to the source for each contact.

I will do this through a Python script that works with the OpenAlex and ORCID APIs to obtain full names, affiliations, positions, and emails where they are public. I will supplement this with parsing department pages and university repositories where emails are not available in the API. I will remove duplicates based on ORCID ID and normalized email, and I will verify the validity of addresses through an SMTP request without sending an email. I will provide the result in one file with all the necessary columns.

It will take 7 days to complete.

Please let me know if there is a priority for specific universities or fields of science, or if a maximally broad sample across all of Ukraine is needed?

The list does not show proposals concealed by the client or freelancer with a Plus profile, as well as proposals violating rules