bettersorted Logo
Shelves with documents – symbolic image for a knowledge base
AI Automation

Building a RAG system in a company: process, costs, common mistakes

AuthorMuhamed Alahmed
Published on
Reading time5 min
In short

A RAG system answers questions based on your own documents: it splits documents into sections, makes them searchable via embeddings, and lets a language model answer only from the retrieved passages. The setup follows six steps; the most common mistakes are poor data quality, missing permissions, and no testing with real questions.

The most important points at a glance

  • RAG answers questions from your own documents – with source references.
  • Six steps: Questions, sources, preparation, indexing, answers, testing & maintenance.
  • Biggest cost driver: Data preparation and permissions concept.
  • Most common mistake: unfiltered drives without permissions.

A RAG system (Retrieval Augmented Generation) answers questions based on your own documents: It breaks documents into short sections, makes them available via Embeddings makes content searchable and allows a language model to answer only from the retrieved passages – with source references. The setup follows six steps. The most common mistakes are poor data quality, missing access rights, and no testing with real questions.

How a RAG system works

The basic idea: Instead of “teaching” a language model your company knowledge, you provide it with the relevant text passages for each question. The model formulates the answer, while the facts come from your documents. More on the term in the glossary under RAG; why this is usually better than retraining is explained in RAG or Fine-Tuning?.

Team planning a project at the whiteboard

The setup in six steps

1. Clarify the objective and questions

Who should ask what? Employees in customer service, in administration, new colleagues? Collect 30 to 50 real questions from day-to-day work. You determine which documents are needed, and they later serve as a test.

2. Select and clean up data sources

Manuals, policies, contracts, product information, FAQs. Sort out outdated versions and duplicates – conflicting documents lead to conflicting answers.

3. Prepare documents

Texts are read from PDFs, Word files, or scans (with OCR for scans) and split into sections. The size of the sections has a strong influence on quality: if they are too small, context is missing; if they are too large, search becomes imprecise.

4. Make searchable

Each section is converted into an embedding and stored in a vector database – together with the source and access rights.

5. Generate answers

When a question is asked, the relevant sections are retrieved and passed to the language model with clear rules: answer only on this basis, cite sources, admit when something is unknown. More on this: AI answers with source references.

6. Test, implement, maintain

Have the collected questions answered, evaluate the responses, and refine them. Then regularly ingest new documents and review open questions.

What a RAG system costs

There are no fixed prices; the main cost drivers are:

  • Data volume and quality: Cleaning up and preparing data is often the biggest effort.
  • Number of sources: Each connection (drive, Nextcloud, DMS, wiki) requires setup.
  • Permissions concept:Who is allowed to see what? The finer the granularity, the more complex it becomes.
  • Model and operation:own GPU server, model service in Germany or cloud – see Host AI models in compliance with the GDPR.
  • Maintenance:continuous ingestion of new documents and analysis.

The most common errors

  • "Everything in": Unfiltered drives with old versions worsen the answers.
  • No permissions: If the system shows personnel files or management documents to everyone, that is a data protection issue.
  • No test: Without real questions, no one knows whether the system helps.
  • No source citation: Without sources, employees cannot verify answers.
  • Only technology, no process: Who maintains documents? Who evaluates open questions?
  • Scanned PDFs without text recognition: content remains invisible.

Off-the-shelf solution or build it yourself?

Many tools already include RAG – for example Nextcloud Context Chat, Open WebUI or Paperless-ngx. For companies with many sources and access requirements, a well-designed solution like our AI document search. How the setup works in detail is described in Set up AI document search.

For those interested in technology

  • Typical setup: document parser with OCR, embedding model, vector database (e.g. Qdrant), language model, orchestration.
  • Hybrid search (embeddings plus classic keyword search) often delivers better results than embeddings alone.
  • A rerankersorts the found sections again by relevance before they are passed to the language model.

bettersorted relies on open, transparent components – operated on servers in Germany and bundled in the automaisa Hub. We provide vendor-neutral advice, are a BAFA-registered consultant and an authorized INQA coach. The Consulting and guided implementation can be subsidized through INQA-Coaching at 80%; the appropriate route is shown by the Funding Check. For an initial, non-binding consultation: Contact.

Example: RAG in a retailer’s customer service

An illustrative scenario: A technical wholesaler has hundreds of data sheets, assembly instructions, and warranty terms. The internal sales team searches daily for answers to customer questions. A RAG system ingests the documents, takes into account which materials are internal only, and answers questions with references to the data sheet and page. After four weeks of testing with 50 real questions, it is released for the internal sales team; open questions are reviewed weekly.

Measure success

  • Share of test questions answered correctly
  • Share of answers with correct source
  • Time saved on typical search tasks
  • Most frequent unanswered questions – indications of missing documents

Costs and funding

The costs of a RAG system consist of three parts: Setup (planning, installation, integration, testing), ongoing operation (server or hosting, updates, monitoring, data backup) and support for employees (training, rules, points of contact). We do not quote fixed prices because the scope and initial situation vary greatly. The software itself does not incur license costs for open-source solutions.

Funding is not for the technology, but for the Consulting and guided implementation: Through INQA-Coaching the federal government covers 80% of coaching costs nationwide (up to €11,520, vouchers until 30.06.2028). A preliminary analysis can be subsidized through the BAFA consulting grant – with 80% in the new federal states, Lüneburg and Trier, otherwise 50%, for applications submitted by 31.12.2026. For investments, some states offer their own programs – see Funding opportunities for companies in Germany.

Frequently asked questions

What is a RAG system?

A system that answers questions based on its own documents: it searches for relevant text passages and has a language model formulate an answer only from those, ideally with a source reference.

How long does it take to build a RAG system?

A pilot with a few well-maintained sources is often possible within a few weeks. The biggest time factor is usually cleaning up and preparing the documents.

Can a RAG system take access rights into account?

Yes, if the permissions are stored during ingestion and checked with every search. This should be planned from the start.

Does a RAG system need its own AI?

No. The language model can run locally, with a German model provider, or in the cloud. For sensitive documents, local processing or processing hosted in Germany is recommended.

As of October 2026.

Portrait of Muhamed Alahmed, founder of bettersorted
About the author

Muhamed Alahmed

With over 10 years’ experience in IT, I develop solutions that not only work from a technical perspective, but also create real added value and open up new possibilities.

More about bettersorted →

More articles

Newsletter

Stay up to date on AI topics

Short updates on AI automation, funding programs and new posts — no spam, unsubscribe anytime.

We use Brevo as our marketing platform. By submitting, you agree that your data will be transferred according to Brevo's privacy policy .

CallEmailContact form