Please select
  • Projects 12
  • Rating 4.9
  • Rating 2 054

Budget: 250 USD Deadline: 7 days

Hello, Andrey! Parsing + auto-posting in Telegram is my specialty, and I have already looked at your sources:
— Douban (movie.douban.com/explore) and Ctrip (you.ctrip.com/travels) load content via JavaScript — I will use Playwright (headless browser), not the "light" BeautifulSoup.
— 36Kr (articles/search/profile) provides data in a more structured way, making parsing easier and faster.

What I will do:
• for each material: title, link, date, author/source, brief text, images (where available);
• auto-tags by source + by topic (from the text);
• deduplication through a database (SQLite) — one material will not be posted twice;
• auto-posting to your Telegram channel via a bot on a schedule;
• configuration (sources/tags/frequency), delays, and logging tailored to the specifics of Chinese sites (frequent requests are cut off) — without bypassing protections, only public data.

  • Projects 5
  • Rating 4.9
  • Rating 756

Budget: 2000 USD Deadline: 7 days

Hello, I worked on a news parser for a media aggregator — collecting articles from 12 sources, ~500 publications/day, auto-posting to Telegram with tags and deduplication.

Regarding your project: some of the mentioned sites (for example, Douban and Ctrip) use dynamic loading via JavaScript — are you open to using Playwright instead of lightweight BeautifulSoup if the site requires it?

I suggest we get in touch; I will provide you with free technical consultation and we can create a development plan + I will tell you about my team!

  • Projects 11
  • Rating 5.0
  • Rating 1 788

Budget: 15000 USD Deadline: 7 days

Good day! We have experience in developing parsers in Python with bypassing protection and integration with the Telegram API. We implement this through the Playwright library for dynamic content and asynchronous message sending to the channel. We will set up stable script operation on the server, taking into account the specifics of Chinese resources.

Maksym Potashov

Maksym Potashov

Winning proposal
6 2
  • Projects 6
  • Rating 3.9
  • Rating 788

Budget: 150 USD Deadline: 7 days

I would start by checking each source separately: Douban, Ctrip, and 36Kr may deliver pages differently, so I will first determine where regular parsing is sufficient and where dynamic loading processing is needed. For the MVP, I would use Python, Playwright/BeautifulSoup, SQLite, Telegram Bot API, and Docker to ensure everything can run smoothly on the server. Duplicates will be stored in the database, and the check frequency and tags will be in the config.

I have worked with Chinese websites, where often the headache is not in the parsing itself, but in the fact that some data is loaded separately or the site may throttle frequent requests. Therefore, I will immediately incorporate delays, logging, and a brief description of the limitations for each site, without bypassing protections.

An MVP for 2-3 sources can be assembled in a few days, and the full version for all sources after checking availability. I will provide a more accurate budget after a quick technical review of the sites.

Please advise, should the publication on Telegram go out immediately after finding the material or go through moderation? Are tags needed only for the source or also for topics within the text? And should the Chinese text be translated/shortened before publication?

Current freelance projects in the category Data Parsing

3:47
21 July
20 July