• Projects 5
  • Rating 5.0
  • Rating 1 321

Budget: 18000 UAH Deadline: 21 days

Hello,

Straight to the stack, since you asked for recommendations rather than instruction-following.

Self-hosted: vLLM to serve the model (Ollama for a lighter footprint), Qdrant as the vector store, LlamaIndex for ingestion and retrieval. Qdrant matters for your isolation requirement: tenant scoping becomes a payload filter inside the query itself, so one client's documents cannot surface in another client's answer even if the application layer has a bug. Ingestion: PyMuPDF for text PDFs with a Tesseract OCR fallback for scans, python-docx and openpyxl for DOCX and XLSX, content-hash deduplication, doc_id plus revision for versioning. FastAPI behind JWT, every answer returning source document, page and chunk so a citation can actually be checked. Docker Compose plus a deployment guide.

About us, honestly. Production Python, FastAPI, Docker and Linux server work delivered on clients' own servers: five completed projects here, average rating 10 out of 10, no negative reviews. We run our own LLM-agent automation in production daily, so agent orchestration and the failure modes of these systems are daily work, not theory. We also have a public record of seven merged pull requests into third-party open-source projects, mostly a Go security tool with 173 stars, each reviewed and accepted by the maintainers. What I will not do is claim a shipped production RAG platform for a client. We build with these frameworks, but I am not going to invent case studies to win a bid.

So de-risk it instead of taking my word. Milestone 1: local LLM running on your server, ingestion for PDF with OCR, DOCX, TXT, XLSX, a Qdrant index with deduplication and versioning, and a FastAPI endpoint answering questions over your real documents with citations. Three weeks, 18000 UAH fixed. You end up with something you can test against your own knowledge base before committing further. Milestone 2 (transcript processing, summaries, action items) and Milestone 3 (permissions, isolation hardening, web interface) I quote after Milestone 1, when real volumes are known. Fixed price per milestone preferred; hourly possible for open-ended research.

  • Projects 16
  • Rating 4.9
  • Rating 4 213

Budget: 700 UAH Deadline: 4 days

Hello, I am in the TOP 10 freelancers on the platform, I have completed such projects before, I will do it quickly, efficiently, and at a low cost.

  • Projects -
  • Rating -
  • Rating 555

Budget: 3600 UAH Deadline: 3 days

What document volume and concurrent user load should the system handle? That decides between self-hosted Qdrant or lighter Chroma, and which local model fits your VPS, CPU-only or with GPU.

I'd deploy Dockerized FastAPI services: LlamaIndex plus Qdrant for RAG, OCR via Tesseract for scanned PDFs, separate agents for transcript summaries and action items, metadata-based access control, source-cited answers.

Milestone-based fits well: first milestone, ingestion and retrieval API, ready in about 3 days for review.

  • Projects 6
  • Rating 5.0
  • Rating 820

Budget: 20000 UAH Deadline: 14 days

Anastasia, setting up a local, air-gapped LLM environment is essential for keeping sensitive business data strictly off third-party APIs. I focus on building resilient RAG pipelines and deploying localized inference engines that prioritize both data privacy and high-accuracy retrieval for your specific documentation workflow. To get the infrastructure right, do you have a preference for the hosting environment or the specific hardware requirements, such as GPU memory capacity, to handle your intended document volume?

The list does not show proposals concealed by the client or freelancer with a Plus profile, as well as proposals violating rules