bettersorted Logo
Office workstation computer – symbolic image for locally operated AI
AI Consultancy

Ollama explained simply: running AI models in your own company

AuthorMuhamed Alahmed
Published on
Reading time5 min
In short

Ollama is a free open-source software (MIT license) that makes it possible to run AI language models such as Mistral, Qwen or gpt-oss on your own computer or server. It downloads models, starts them, and makes them available via an interface. For companies, it is an easy way to get started with AI without the cloud.

The most important points at a glance

  • Ollama is free open-source software (MIT) for running AI models.
  • It downloads, starts, and serves models – local, without cloud.
  • For teams, it is combined with an interface such as Open WebUI.
  • Most important hardware metric: graphics memory (VRAM).

Ollama is free open-source software (MIT license) that lets you run AI language models such as Mistral, Qwen, or gpt-oss on your own computer or server operate. Ollama downloads the desired model, starts it, and makes it available via an interface that other programs can use. For companies, it is the easiest way to try out and use AI, without sending data to a cloud.

What Ollama does exactly

You can think of Ollama as the “engine room” for AI models. It takes care of three things:

  • Getting models:Open models can be downloaded from an extensive library with a single command – in suitable, often smaller variants.
  • Run models: Ollama automatically uses available graphics cards and falls back to the processor if necessary.
  • Provide models: Via an interface, chat front ends, automations, or custom applications can access the model – also in OpenAI-compatible format.

Ollama itself does not have an elaborate user interface for teams. For that, it is combined with tools such as Open WebUI, which offers a ChatGPT-like interface with user management.

IT specialist at a server rack

What companies use Ollama for

  • Internal AI assistant: Draft, summarize, and translate texts – without content leaving the company.
  • Working with your own documents: In combination with a knowledge base via RAG the model answers questions about your documents.
  • Automation: workflows, for example with n8n, use the model to sort emails or extract information from documents.
  • Document processing: tools such as Paperless-ngx or NextcloudYou can connect a local model via Ollama.
  • Try it out: Test several models side by side before deciding.

What Ollama is not

  • Not a model: Ollama is the runtime environment; the quality of the responses depends on the model selected.
  • No complete solution: User management, permissions, logs, and data backup are added through additional components.
  • Not ideal for every workload: With many concurrent users, Ollama reaches its limits; in that case, specialized servers such as vLLM are used.

What hardware do I need?

Small models already run on a well-equipped workstation, medium-sized models need a graphics card with sufficient memory, and large models require a GPU server. The key factor is above all the Graphics memory (VRAM). A detailed assessment is provided by What hardware does a dedicated AI need?.

Data protection: the major advantage

If Ollama runs on your own hardware or with a hosting provider in Germany, inputs and documents do not leave your infrastructure. There is no third party that stores inputs or uses them for training. Still, the following applies: even a local AI processes personal data – you must organize access rights, logs, and deletion rules yourself. More on this: Host AI models in compliance with the GDPR.

How companies proceed

  1. Define the objective: Which tasks should the AI take over?
  2. Pilot: Ollama with two or three models on a test machine, evaluation with real tasks.
  3. Interface and permissions: Open WebUI or another interface with login via Single Sign-On.
  4. Operation: Servers in Germany, updates, backups, monitoring – in-house or through a service provider.
  5. Rules: An AI policy that defines what AI may be used for.

For technology enthusiasts

  • Ollama runs on Windows, macOS and Linux and can also be operated as a container.
  • Models are stored in quantized form (usually 4-bit) loaded – this saves memory with only a minor loss in quality.
  • By default, the interface is only accessible locally; for network access, a secured gateway should be placed in front of it.

bettersorted relies on open, inspectable components for its customers – operated on servers in Germany and bundled in the automaisa Hub. We provide manufacturer-neutral consulting, are a BAFA-registered consultant and an authorized INQA coach. The consulting and guided implementation can be subsidized through INQA-Coaching at 80%; the appropriate path is shown by the Funding Check. For an initial, non-binding consultation: Contact.

Example: Craft business with quotation templates

An illustrative scenario: An electrical contractor with 25 employees wants to use AI to prepare quotation texts, customer letters, and summaries of site reports. A rented server with a graphics card runs Ollama with a medium-sized model, and Open WebUI runs on it with login access for office staff and site management. Customer data remains in Germany; after one month, the templates are refined and the most common tasks are stored as templates.

Ollama compared to other runtime environments

  • Ollama: very easy, ideal for getting started and for small teams.
  • vLLM: for many concurrent users and high load, more complex to set up.
  • llama.cpp:the technical foundation of many tools, very flexible, more suited to specialists.
  • Model services in Germany:no in-house operation, usage-based billing – see Host AI models in compliance with the GDPR.

Common mistakes

  • Make Ollama reachable on the network without protection.
  • Choose a model that is too large for the available graphics memory.
  • No rules for what the assistant may be used for.
  • Test only one model instead of comparing two to three candidates.

Costs and funding

The costs of using Ollama in the company consist of three parts: Setup (planning, installation, integration, testing), ongoing operation (server or hosting, updates, monitoring, data backup) and support for employees (training, rules, contact persons). We do not quote fixed prices because the scope and starting point vary greatly. The software itself incurs no license costs for open-source solutions.

Funding is not for the technology, but for the consulting and guided implementation: Through INQA-Coaching the federal government covers 80% of coaching costs nationwide (up to €11,520, vouchers until 30/06/2028). A preliminary analysis can be funded through the BAFA consulting grant subsidize – with 80% in the new federal states, Lüneburg and Trier, otherwise 50%, for applications submitted by 31.12.2026. For investments, some states offer their own programs – see Funding opportunities for companies in Germany.

Frequently asked questions

Is Ollama free of charge?

Yes. Ollama is open-source software under the MIT license and can be used free of charge, including commercially. Costs arise for hardware or hosting and operation.

Is Ollama safe for company data?

Ollama processes data locally on the computer on which it runs. For use in a company, access, network, and permissions still need to be properly secured — for example, through an interface with login.

Which models can I use with Ollama?

Many open models, including Mistral, Qwen, gpt-oss, DeepSeek, and Gemma. Check the license for each one — Llama 4 is not licensed for companies based in the EU.

What is the difference between Ollama and Open WebUI?

Ollama runs the models, Open WebUI is the user interface for people. In practice, both are usually used together.

As of October 2026. Models, versions and licenses change quickly – check the current license with the provider before making a decision.

Portrait of Muhamed Alahmed, founder of bettersorted
About the author

Muhamed Alahmed

With over 10 years’ experience in IT, I develop solutions that not only work from a technical perspective, but also create real added value and open up new possibilities.

More about bettersorted →

More articles

Newsletter

Stay up to date on AI topics

Short updates on AI automation, funding programs and new posts — no spam, unsubscribe anytime.

We use Brevo as our marketing platform. By submitting, you agree that your data will be transferred according to Brevo's privacy policy .

CallEmailContact form