Skip to content
Local AI

Local AI Data stays on your hardware

With local AI you process confidential data directly on your own hardware. Contracts, customer data, personnel files or internal documents do not leave your company. I use local AI where cloud solutions are off the table for data protection, compliance or organisational reasons.

As a technical base I use Ollama among other tools. That lets open-weight models run on a Mac, on Windows or on a Linux server. Prompts, documents and answers stay on your own infrastructure.

Orientation

What local AI with Ollama actually is

Ollama simplifies running modern open-weight models on your own hardware. What used to need many tools and a lot of configuration can now be stood up locally in a few steps. At the same time Ollama provides an API that is straightforward to wire into existing apps or automations.

For many typical company tasks, current models on modern business hardware are already enough. Which model size makes sense depends on the hardware you have and the speed you need. Not every task needs the strongest model from the cloud.

Ollama is now the most common standard for local AI, but not the only option. Depending on the use case, LM Studio or llama.cpp can be the better choice.

The problem

Confidential data does not belong in public AI services

Contracts, bank statements, personnel files or customer data should not be pasted into public AI services without review. In many companies that raises questions about data protection, compliance or internal policy.

Local AI makes it possible to use language models for exactly those tasks without sensitive content leaving the company.

01

Data protection and compliance

Even with a DPA and an EU region, legal questions remain with US vendors. For many companies, public bodies or regulated industries, local processing is therefore the only option on the table.

02

Hardware decides the speed

Local AI only works well when model and hardware fit. A model that is too large for the graphics memory leads to long wait times and poor day-to-day adoption.

03

Single seat or team setup

For individual staff a capable workstation is often enough. As soon as several people need the same AI at once, a central server is the better fit.

Setup

Three ways to run local AI in a company

Not every company needs its own GPU server. The right setup follows the number of users, the data risk and the use case.

Workstation

Ideal for individual staff, developers or confidential documents that must not leave the machine.

Office server

A central Linux server provides the AI on the internal network. Apps such as n8n or your own software talk to it through a shared interface.

Own instance in a datacentre

The same software runs on rented hardware in a European datacentre. You skip hardware on site, while still controlling the infrastructure.

Mac

Apple Silicon uses shared memory for CPU and GPU. That lets many mid-size models run usefully on ordinary machines.

Windows and Linux

Here the available graphics memory of the NVIDIA card mainly decides which models can run at a useful speed.

The solution

Local AI fitted to your company

I pick model, infrastructure and hardware to fit your use case. That may be Qwen or Llama with Ollama, LM Studio or llama.cpp. The technical base stays flexible and can later be extended or swapped.

Local AI does not automatically mean full isolation. Downloading a model hits a model library once. Optional cloud features should stay off. A local AI server must also not be reachable from the internet without protection.

What local AI does especially well

  • Summarise and reply to emails
  • Search and summarise internal documents
  • Build knowledge bases with RAG
  • Support invoices and incoming payments
  • Categorise and structure text
  • Help developers with individual coding tasks
  • Process confidential company data locally

Local models do not automatically replace the strongest cloud models. For complex analysis, very large document sets or demanding agent systems, a cloud setup often stays ahead. In many companies a hybrid approach is therefore the sensible path.

Tools

Ollama, LM Studio or llama.cpp

Ollama fits most companies as the default for local AI. LM Studio offers a usable interface without a terminal. llama.cpp is aimed at technical users who need maximum control or speed.

Which setup is used depends on the company requirements, not on a particular piece of software.

What you should not force locally

  • Very many concurrent users on a single server
  • Tasks where small local models demonstrably hit their limits
  • Oversized models on unsuitable hardware
  • Permanently loaded office PCs as a stand-in for proper infrastructure
In practice

Where local AI is used day to day

Local AI fits especially well for processes with sensitive data. The model helps where classic rules or scripts are no longer enough.

01

Match incoming payments

If the payment reference has typos or mismatched company names, the local model helps with matching. Unclear cases are still reviewed by staff.

02

Internal knowledge search

Approved documents can be searched through a RAG setup. Answers are based only on the company data you provide.

03

Email and classification

Inbound messages can be categorised, prioritised or prepared as a draft.

04

Hybrid AI setups

Sensitive data is processed locally. For especially demanding tasks an enterprise cloud setup can be attached if needed.

Questions

Common questions about local AI

Does the data really stay in the company?

Yes, as long as only local models are used. Prompts, documents and answers stay on your own hardware. Exceptions are the one-time model download and cloud features that are deliberately switched on.

Is local AI GDPR-compliant?

Local AI makes it much easier to process sensitive information in a data-protection-friendly way. Whether a concrete setup meets every legal requirement still depends on the use case.

What hardware is needed?

The right hardware follows model size and use case. Modern business machines are already enough for many tasks. More demanding work benefits from dedicated graphics cards or a central server.

Can ChatGPT be fully replaced?

For many internal tasks, yes. For complex analysis or especially capable models, a cloud setup often remains the better choice.

Can existing company knowledge be used?

Yes. With RAG you can search documents, wikis or SharePoint content without sending it to public AI services.

Orientation

When cloud AI is the better fit

Local AI gives maximum control over confidential data. For especially complex tasks, many concurrent users or very large models, a cloud setup is often more economical.

The other option

Advantages of cloud AI

Strong models, flexible scaling and no hardware of your own. In return, data protection, the contract and where data is stored should be checked carefully.

To cloud AI
Starting point

Check the use case first

Not every process needs AI. I first check data risk, benefit and economics. After that it is clear whether a local setup, a cloud setup or a mix of both makes sense.

To the AI check

In 15 minutes, see whether local AI makes sense for your company

In the first conversation I clarify with you:

  • whether local AI fits your use case
  • what hardware you already have
  • which model size makes sense
  • whether a cloud setup would be more economical
  • which next steps follow from that

After that you have a realistic view of which setup offers the most value for your company.