Website parser for mouser.pl
- Regular parsing is needed on a schedule
- Authorization is not required
- Translation is not needed
- Markup is also not needed
- The site has a lot of products, over 1 million, so the parser needs to be divided into several separate groups; I will not parse all groups, only some. First of all, we need to start with the group Rozwiązania typu embedded / Embedded solutions, which has 42,000 products
- The parser should be installed on a VPN server and run automatically on a schedule using CRON
- The result of the work - several YML (XML) files in PROM format - the description of the format is here
https://support.prom.ua/hc/uk/articles/360004963538-%D0%86%D0%BC%D0%BF%D0%BE%D1%80%D1%82-%D1%87%D0%B5%D1%80%D0%B5%D0%B7-YML-%D1%84%D0%BE%D1%80%D0%BC%D0%B0%D1%82-%D1%84%D0%B0%D0%B9%D0%BB%D1%83
- A filter for groups and subgroups is needed, a blacklist, as well as a price filter from and to, in stock

- Do not parse products: Stock levels 0, without photos, with AR, EAR icons



Data collection
example of a product https://www.mouser.pl/ProductDetail/Taoglas/TG.66.AN13?qs=sGAEpiMZZMu3sxpa5v1qrmhEv3glFT0EBOUeqiKfExg%3D


Description of 2 options
1. More detailed information


With HTML tags for photos and videos, but remove links
2 If there is no detailed information
example https://www.mouser.pl/ProductDetail/Innodisk/EMUC-B202-W2?qs=sGAEpiMZZMsGx6xItMI%252BGHM48Cf3t62hm4W8FdDvs4Y%3D
then duplicate the title in the description