Budget: 1000 UAH Deadline: 1 day
Good day. I am ready to take a look at how the parser works, and after that, I will name the price.
The parser is written in Python. The library implements data collection for readability.
Documentation: https://pypi.org/project/readability/
The developer recorded the technical specifications in voice, explaining the problem. The specifications are in the archive. 1 voice note for Readability and 2 voice notes for Bootstrap. Also, when I accept you into the project, I will pass on any of your questions to the developer.
He managed to do it only this way. For beginners, this is probably not feasible. Therefore, I am reaching out to professionals.
Regarding Bootstrap. He also tried to implement it, but Bootstrap produced worse results compared to Readability. There was a lot of duplicated content that was the same. And it picked up extra unnecessary dirty code.
About the code itself: the code is written in Python. Requests to the server are made through aiohttp,
because the project is asynchronous, meaning requests are sent to the server in parallel, not sequentially.
The build is done using the PyInstaller library. I run the .exe program, and the command line opens. The parser itself opens in the browser, locally at address 127: and so on.
To evaluate the code and the cost of work. And you do not write a number out of thin air. I understand you. You will write a conditional one. Therefore, a convenient option. Connect to my PC. Look at the code. Understand that you can improve the parsing results and solve the task so that it picks not only text but also images from the websites. Then you will update your bid under the project, and I will accept you into the project. I will allocate a reserve of funds. And only then! Because! If you do not look at the code and write any bid. What will be the outcome? My time wasted and funds? And a negative review for you? I think you don't need that. I think we clarified this. Now, such a result for example from 10 websites. Out of 10 websites, it only picks text from 5 websites, and from the other 5 websites, it picks text + images. It picks text from all 10 websites. I think the logic is clear. What is needed is for it to also pick images just like text from all websites.
I don't care how to implement it through Readability or Bootstrap. What matters to me is that the parser picks data more accurately. Through Readability, it picks text from each site, but not images from each. Therefore, the task was to improve it or cross it with another library, algorithm, technology. That would pick images. And it would pick text.
Or to do it entirely through Bootstrap. But only so that it picks both text + images from all websites. In short, it should work on Bootstrap no worse than on Readability.
I can provide access through Anydesk, I can compile and build it into bild.exe myself. You just need to log into my PC, evaluate the code. And see if you can make changes in my code. On bs4. If you think that this will improve data collection and solve my problem, then no questions asked. If we test together and see that your technology is better, I will immediately choose you for the project. I will allocate a reserve of funds, you will make changes to the code. We will test. If the results are better, I will accept the project.
Budget: 16000 UAH Deadline: 1 day
Hello,
I am ready to take on your Python parser project for data collection using the Reability library. I have experience in developing code in Python and using aiohttp for asynchronous requests. Bundled application launch via PyInstaller is also in my arsenal.
To evaluate the code and develop a strategy for collecting both text and images from websites, I invite you to connect to my PC via anydesk. Upon a deeper review of the code and testing, we can make the necessary changes and improvements to achieve the desired result.
My hourly rate is $16. I look forward to your response for further collaboration.
Best regards,
Maxim
Доброго дня Александр
Вашу програму можна покращити, але це не буде саме те, що Ви хочете.
Розбирати правильно абсолютно будь який сайт неможливо, або близько до цього.
Як мінімум -- на данний час.
В те щоб зробити readability вкладено багато грошей і років часу.
Якщо у Вас є якийсь перелік сайті(лінків) які Ви регулярно скрейпите -- то надішліть мені. Я подивлюсь який відсоток вийде покращити.
Зараз я трохи зайнятий і не зможу відповідати миттєво
Web Programming 25 proposals 30 July
Python 39 proposals 30 July
Bot Development 31 proposals 30 July
Web Programming 66 proposals 29 July
Data Processing 33 proposals 28 July