About the project:
We are developing an AI platform for automating document flow. The system is already in production but has critical issues with the quality of document search and analysis. We are looking for an experienced architect-engineer who will not only identify the problems but also directly help to solve them.
Current situation:
A working microservices system based on LLM with document search
Low relevance of search results
Issues with processing large/complex documents
RAG-pipeline needs redesign and optimization
The system processes PDF/DOCX/XLSX + OCR, indexing in Weaviate
What needs to be done:
Audit:
Analyze the current architecture of the RAG system (Python FastAPI + LangChain)
Evaluate the effectiveness of the document processing pipeline
Check the quality of work with Weaviate (indexing, retrieval, chunking)
Identify critical issues and their causes
Determine priorities for refactoring
Implementation:
Redesign the chunking strategy and document processing pipeline
Optimize the embedding pipeline and work with Weaviate
Improve RAG pipelines (indexing/citation)
Optimize prompt engineering and retrieval logic
Implement metrics and evaluation for search quality
Set up an optimal flow for processing different types of documents
Integrate improvements with the existing Chat API
Assist with deployment in production
Knowledge transfer:
Document architectural decisions and changes
Conduct sessions with the dev team
Set up monitoring and evaluation metrics
Expected result:
A working optimized RAG architecture
Increased accuracy and relevance of search (measurable metrics)
Faster and higher quality document processing
Documentation and best practices for the team
The ability to maintain and develop the system independently
Required skills:
Must have:
Practical experience in building production RAG systems
Hands-on experience with Weaviate (or willingness to quickly learn)
Experience with vector databases in general (Pinecone, Qdrant, Milvus)
Deep knowledge of embedding models and semantic search
Experience with LangChain (Python) — this is our main framework
Strong Python skills (FastAPI will be a plus)
Experience with document parsing/OCR (PDF, DOCX, XLSX, images)
Understanding of chunking strategies, hybrid search, reranking
Nice to have:
Experience with evaluation frameworks for RAG systems (RAGAS, LangSmith, etc.)
Knowledge of cost optimization methods in LLM applications
Experience with microservices architecture
Understanding of Node.js/TypeScript for integration with Chat API
Knowledge of PostgreSQL for working with metadata
Our tech stack:
Backend:
Python (FastAPI) — documents, RAG pipelines, LangChain
Node.js (Fastify) — core, Auth, Chat API orchestration
Go — high-load I/O tasks
Data & Storage:
PostgreSQL (Drizzle ORM) — metadata, users, chats
Weaviate — vector search, semantic indexing
MinIO (S3-compatible) — files, artifacts
Redis — cache, queues (BullMQ), sessions
AI Stack:
LangChain (Python)
Document processing: PDF/DOCX/XLSX parsing, OCR
RAG pipelines with indexing and citation
Frontend:
React + TypeScript + Vite
REST API + WebSocket
Collaboration format:
Close collaboration with the dev team (Python/Node.js developers)
About the team:
Early-stage startup with a working MVP in production. Microservices architecture, an experienced technical team ready to actively work with the architect and quickly implement changes.