Budget: 9999 UAH Deadline: 6 days
Hello, Konstantin! Your task requires a systematic approach to processing unstructured data. I have experience working with PDF analytics and developing scalable Python projects.
My implementation plan according to your requirements:
Architecture: I will build a modular structure (OOP), where each type of statement is a separate plugin-module. This will allow for easy addition of new banks without changing the core of the system.
Hybrid parsing: I will use pdfplumber for instant text extraction and EasyOCR/Tesseract for graphical elements (stamps, handwritten dates). This will ensure a speed of 1-3 seconds per file.
Normalization: I will create a universal data schema (Transaction Model). In the output, you will receive clean JSON or DataFrame with validated fields (date, amount, purpose, balance).
Training: I will train the logic on your 3 types of statements, ensuring resilience to layout shifts and specific bank encodings.
No cost: I will use exclusively open-source solutions without reliance on paid cloud APIs.
I am ready to discuss the structure of the output template and start developing the prototype.
Best regards,
Victor