bettersorted Logo
Graphics card in a computer – symbolic image for AI hardware
AI Consultancy

What hardware does a custom AI need? GPU servers for mid-sized businesses

AuthorMuhamed Alahmed
Published on
Reading time6 min
In short

For a custom AI, graphics memory (VRAM) is the key factor. As a rule of thumb, a model reduced to 4 bits needs about 0.6 to 0.7 GB of VRAM per billion parameters, plus some reserve. Small models run on a single graphics card, while large ones need GPU servers. Instead of buying, SMEs can rent GPU servers in Germany.

The most important points at a glance

  • The decisive factor is the graphics memory (VRAM), not the processor.
  • Rule of thumb: 0.6–0.7 GB VRAM per billion parameters (4 bits) plus reserve.
  • Three stages: Test computer, a GPU server, multiple GPUs or model service.
  • For many SMEs, Renting in Germanythe better entry point.

Whether a company's own AI runs in-house depends primarily on one value: the graphics memory (VRAM) of the graphics card. As a rule of thumb, a language model reduced to 4 bits requires around 0.6 to 0.7 GB of VRAM per billion parameters – plus reserve for longer texts and multiple users. Small and medium-sized models run on one graphics card, large ones require GPU servers. Instead of buying hardware, SMEs can also rented in German data centers.

Why the graphics card, of all things?

Language models work with huge numerical tables. Graphics cards are specialized for exactly this kind of computation and are much faster than standard processors. For the model to respond quickly, it must fit entirely into the graphics card memory. If there is not enough memory, it either becomes very slow or does not run at all.

Technician working on server hardware

Rules of thumb for memory requirements

The following values are rough guideline values for models with 4-bit quantization and one user; depending on the model and settings, they may vary:

  • Small models (around 8 billion parameters): about 6–8 GB VRAM – fits on many current graphics cards.
  • Medium-sized models (around 20–30 billion): about 16–24 GB VRAM – a powerful single card.
  • Large models (around 70 billion): about 40–48 GB VRAM – a professional card or two cards.
  • Very large models (100 billion and more): several professional cards or specialized servers.

In addition, storage is needed for the “context” — that is, the texts the model should keep in view at the same time — and for additional concurrent users. Always plan for reserve capacity.

Three typical expansion stages

Stage 1: Testing on existing hardware

A workstation with a good graphics card or a current Mac with plenty of shared memory is enough to test small models with Ollama. Ideal for a pilot, not for team operations.

Stage 2: A GPU server for the team

A server with a powerful graphics card runs a medium-sized model for a small team — for example as an internal assistant or for document search. German providers rent out such servers; Hetzner, for example, offers a dedicated server with an NVIDIA RTX 4000 SFF Ada with 20 GB of graphics memory.

Stage 3: Multiple graphics cards or a model service

For large models or many concurrent users, multiple professional cards are needed. In this case, it is often worth comparing with a model service in Germany, where you do not operate a server but only pay for usage — see Host AI models in compliance with GDPR.

Buy or rent?

  • Buying is worthwhile with consistently high utilization, an available server room, and in-house IT staff.
  • Renting is flexible, with no initial investment, professional power supply, cooling, and network connectivity – and can be scaled up or down as needed.
  • Power and heat don't forget: GPU servers consume significantly more energy than standard office equipment and require cooling.
  • Replacement and maintenance: If the only graphics card fails, the AI stops. Rental servers are replaced by the provider.

Typical errors

  • Too little graphics memory: The model does not fit or runs only at a much slower speed.
  • Planned for a single user only: Multiple simultaneous requests require additional capacity.
  • Model too large: Often, a medium-sized model with a solid knowledge base via RAG is better than a huge one without it.
  • Underestimated in operations:Updates, data backup, and monitoring take time or require a service provider.

For technology enthusiasts

  • For multi-user operation, inference servers such as vLLM, which bundle requests and manage memory efficiently, are suitable.
  • With MoE models all parameters must be loaded into memory, even though only part of them is used for each response.
  • Data center GPUs offer more memory and error correction (ECC), but are significantly more expensive than consumer cards.

bettersorted relies on open, inspectable components – operated on servers in Germany and bundled in the automaisa Hub. We provide vendor-neutral advice, are a BAFA-registered consultant and an authorized INQA coach. The consulting and guided implementation can be funded through the INQA-Coachingfund with 80%; the appropriate path is shown by the Funding Check. For a non-binding initial consultation: Contact.

Example: calculation for a team of 15 people

An illustrative scenario: 15 employees use an internal AI assistant, usually not at the same time; peaks are three to four parallel requests. A medium-sized model is desired (around 30 billion parameters, 4-bit). As a rule of thumb, this requires about 18–21 GB of graphics memory for the model plus reserve for context and parallel requests – a 24 GB card is tight, while 32 to 48 GB is comfortable. Alternatively, a somewhat smaller model can be combined with good document retrieval and run on a 20 GB card.

Questions to ask the host before renting

  1. Which graphics card, how much graphics memory, how much RAM?
  2. Data center location and data processing agreement?
  3. How quickly is defective hardware replaced?
  4. Are there availability commitments and data backups?
  5. Can the server be expanded later or can a second one be added?

Energy and sustainability

Under load, GPU servers consume several times as much power as an office PC. A medium-sized model that meets the requirements is not only cheaper, but also more energy-efficient than a large one. Data centers with green electricity and efficient cooling are an additional selection criterion.

Costs and funding

The costs of dedicated AI hardware consist of three parts: Setup (planning, installation, integration, testing), ongoing operation (server or hosting, updates, monitoring, data backup) and support for employees (training, rules, contacts). We do not quote fixed prices because the scope and starting point vary greatly. The software itself does not incur license costs for open-source solutions.

Funding is not for the technology, but for the Consulting and guided implementation: Through INQA-Coaching the federal government covers 80% of coaching costs nationwide (up to €11,520, vouchers until 30.06.2028). A preliminary analysis can be subsidized through the BAFA consulting funding – with 80% in the new federal states, Lüneburg and Trier, otherwise 50%, for applications submitted by 31.12.2026. For investments, some states offer their own programs – see Funding opportunities for companies in Germany.

Frequently asked questions

How much graphics memory does a custom AI need?

As a rule of thumb, about 0.6 to 0.7 GB per billion parameters with 4-bit quantization, plus headroom. A model with around 8 billion parameters therefore needs about 6 to 8 GB, and one with 70 billion about 40 to 48 GB.

Is a standard office PC sufficient for AI?

For small models to try out, yes, sometimes. For team use, you need a computer or server with a powerful graphics card.

Should an SME buy or rent a GPU server?

For most SMEs, renting in a German data center is the better starting point: no upfront investment, professional operation, and flexible scaling. Buying is mainly worthwhile with consistently high utilization and in-house IT.

Can AI run without a graphics card?

Small models can also run on the processor, but much more slowly. For everyday use in a team, a graphics card is practically a requirement.

As of October 2026.

Portrait of Muhamed Alahmed, founder of bettersorted
About the author

Muhamed Alahmed

With over 10 years’ experience in IT, I develop solutions that not only work from a technical perspective, but also create real added value and open up new possibilities.

More about bettersorted →

More articles

Newsletter

Stay up to date on AI topics

Short updates on AI automation, funding programs and new posts — no spam, unsubscribe anytime.

We use Brevo as our marketing platform. By submitting, you agree that your data will be transferred according to Brevo's privacy policy .

CallEmailContact form