
Vector databases explained simply – using Qdrant as an example
A vector database stores texts, images or other content as sequences of numbers (embeddings) and finds the entries that are most similar in content to a question – even without the same words. It is the memory of RAG systems. Qdrant is a widely used open-source vector database under the Apache 2.0 license from a Berlin-based company.
The key points in brief
- Vector databases search for meaning instead of words.
- They are the memory of RAG systems.
- Qdrant: Open Source (Apache 2.0), provider from Berlin, self-hostable.
- Data protection: vector database how to protect the documents themselves.
A vector database stores content – usually text passages – as sequences of numbers (so-called embeddings) and finds the most semantically similar entries for a question, even if they contain completely different words. It is the “memory” of AI systems that work with their own documents. Qdrant is a widely used open-source vector database under the Apache 2.0 license, developed by Qdrant Solutions GmbH from Berlin.
Why conventional databases are not enough
Traditional databases and search functions look for identical words. If someone searches for “vacation request,” they won’t find the “absence notification.” Vector databases search for meaning: They compare how close the sequences of numbers in two texts are. Content with similar meaning lies close together in this “semantic space.”

How a vector database works – in four steps
- Ingestion: Documents are split into sections.
- Convert: An embedding model turns each section into a sequence of numbers.
- Store: The vector database stores the sequence of numbers together with the text, source, and additional information — such as department or access rights.
- Search: The question is also converted; the database returns the sections with the closest meaning — in milliseconds, even with millions of entries.
These sections are then passed to a language model so it can formulate a response from them – the principle of RAG, explained in detail in Building a RAG system.
Qdrant in profile
- Open Source: Apache-2.0 license, with no functional restrictions when self-hosted.
- European provider: Qdrant Solutions GmbH, based in Berlin.
- Self-hostable: on your own server or with a host in Germany; in addition, the manufacturer offers a cloud service.
- Filter: Searches can be combined with conditions – for example, “only documents from the Administration department” or “only what this person is allowed to see.”
- Powerful: written in the Rust programming language, designed for large volumes of data.
Examples of use in the company
- Document search:Questions about manuals, contracts and policies – see AI document search.
- Chatbots: Answers from your own knowledge base – see AI chatbot.
- Find similar cases: previous cases, complaints, or quotations with similar content.
- Detect duplicates: nearly identical documents or customer records.
Data protection in vector databases
Sequences of numbers can also allow conclusions to be drawn about content, and the database usually stores the original text as well. It therefore needs to be protected just like the documents themselves: operation in Germany, access restrictions, encryption, and a deletion concept. If a document is deleted, its entries in the vector database must also disappear.
Does every company need its own vector database?
No. Many tools come with a built-in solution. A standalone vector database such as Qdrant is worthwhile when multiple applications access the same knowledge base, large volumes of data need to be searched, or fine-grained permissions are required.
For technically interested readers
- Qdrant offers REST and gRPC interfaces and runs as a single container or in a cluster.
- In addition to dense vectors, it supports Sparse vectors for hybrid search (meaning plus keywords).
- Alternatives include pgvector (extension for PostgreSQL), Weaviate or Milvus.
bettersorted relies on open, inspectable components – operated on servers in Germany and bundled in the automaisa Hub. We provide manufacturer-neutral advice, are a BAFA-registered consultant and an authorized INQA coach. The consulting and guided implementation can be subsidized through INQA-Coaching at 80%; the appropriate route is shown by the Funding Check. For an non-binding initial consultation: Contact.
Example: find similar complaints
An illustrative scenario: A manufacturer of components receives complaints in free text every week. A vector database stores all previous cases together with the solution. When a new complaint comes in, the system finds the three most similar earlier cases – even if customers describe the problem differently. The service team can immediately see which solution worked at the time.
When you need your own vector database
- Multiple applications are to access the same knowledge base.
- It involves large document collections.
- Responses must be filtered by permissions, departments, or tenants.
- You want to be able to replace the embedding model or language model independently.
Costs and funding
For a knowledge base with a vector database, open-source tools do not incur license costs. Costs arise for setup (planning, installation, integration, testing), operation (server or hosting in Germany, updates, monitoring, data backup) and support (training, rules, contact persons). We do not quote fixed prices because scope and starting point vary greatly.
Eligible for funding is the consulting and guided implementation: The INQA-Coaching covers 80% of coaching costs nationwide (up to €11,520, vouchers until 30.06.2028); a preliminary analysis is subsidized by the BAFA consulting subsidy with 80% in the new federal states, Lüneburg and Trier, otherwise 50% – for applications submitted by 31.12.2026.
Frequently asked questions
What is a vector database?
A database that stores content as sequences of numbers (embeddings) and finds the entries most similar in content to a query – even without the same words.
What is a vector database used for?
Especially for AI applications with your own documents (RAG), for semantic search, chatbots, finding similar cases, and detecting duplicates.
Is Qdrant free of charge?
Yes, Qdrant is open source under Apache 2.0 and can be self-hosted free of charge. Costs arise for servers and operations, or for the optional cloud service provided by the vendor.
Is a vector database relevant under GDPR?
Yes. It usually contains original texts and derived data from documents. It must be protected, deleted, and documented just like the documents themselves.
As of October 2026.

Muhamed Alahmed
With over 10 years’ experience in IT, I develop solutions that not only work from a technical perspective, but also create real added value and open up new possibilities.
More about bettersorted →More articles

Embeddings Explained Simply: How AI Searches for Meaning Instead of Words
Learn what embeddings are, how semantic search works, and why they matter for companies—explained simply, without math. Read now.

Building a RAG system in a company: process, costs, common mistakes
This is how companies build a RAG system: six steps, cost drivers, common mistakes, and data protection – clearly explained for decision-makers. Read now.

RAG or Fine-Tuning? How AI really uses your company knowledge
RAG or Fine-Tuning: What is the difference, what is suitable for company knowledge, what costs what? Clearly explained with decision support. Read now.
Stay up to date on AI topics
Short updates on AI automation, funding programs and new posts — no spam, unsubscribe anytime.