Private AI hosting

AI that never leaves your network.

An open model runs on a machine you own, or on one that runs for you inside Switzerland. It has no route to the internet, so nothing can travel out: not a question, not a file, not a name from your client list.

  • No route to the outside
  • No provider trains on your content
  • Fixed price after the consultation
When your own server is the answer

Four situations where going through a provider is simply not on the table.

Work under professional secrecy

Law firms, fiduciaries, medical practices. What somebody tells you belongs to them and to nobody else. Sending it to another organisation’s server is not a matter of convenience here, it is a problem.

Data that stays in the country

A contract, an agreement with your clients or a rule in your sector requires the data to stay in Switzerland. Then the deciding factor is where the machine stands, not what it costs.

An internal policy against the cloud

There is a directive in the house and it says no to the cloud. Rather than fight for an exception, we build the AI so that the question never comes up at all.

Genuinely sensitive material

Personnel files, contracts, medical histories, salary data. The very documents where AI takes the most work off your hands are the ones least allowed to travel.

What is included

From the model on your own machine to the log that shows who asked what.

An open model on your machine

Llama, Mistral, Qwen or Gemma running on a computer in your server room or in a Swiss data centre. Which model it turns out to be follows your task, not the biggest name on the market.

An assistant on your internal material

Instructions, contracts, manuals, case files. The model searches your own documents and writes the answer from them, with a pointer to the passage it came from.

No connection to anything outside

There is no call out to a provider, because none is needed. The machine is not permitted onto the internet at all, and that can be verified from inside your own network.

Access control and a full log

Who may use the assistant, and which material it shows them, hangs off your existing accounts. Every question and every answer is recorded, so who asked what can still be traced later.

Change the model when you choose

When a better open model appears, we swap it in. You pick the moment, not a provider switching a version off. Your documents and your settings stay exactly where they are.

Measured performance and consumption

Response times, graphics card load and power draw are all readable. That way you see whether the machine still has room before people in the house start complaining.

What it actually takes

A machine with a graphics card, and no more model than the job needs

The usual worry is the outlay. It tends to come out smaller than feared, because the size of the model follows the task rather than the other way round.

The hardware

A computer with a suitable graphics card, either in your own server room or rented in a Swiss data centre. Rented means no purchase and no server room, and the data still stays in the country.

The size of the model

A model does not have to do everything, only your work. Searching contracts and writing summaries calls for a different calibre than generating code. We measure that against your real material before any hardware is ordered.

Trial first, purchase second

The first run happens on borrowed hardware. If it turns out that your own server is not worth it in your case, we say so plainly, and you have bought nothing.

Which model, and why

A mid-sized model with your documents beats the largest one without them

Open models such as Llama, Mistral, Qwen or Gemma are good enough today for the work most businesses actually have. What decides the outcome is rarely the size of the model, but what it gets to see before it answers.

Why the context does the work

A very large model knows a great deal about the world and nothing about the instruction you issued two years ago. A mid-sized model that searches your documents before every answer does better on precisely those questions.

How we choose

We take your own questions, let two or three models answer them and lay the results side by side. The decision follows that comparison, not a league table somewhere on the internet.

Open does not mean the same everywhere

The licence terms of open models differ, particularly on commercial use. We read them through with you before the model is installed in your house.

What happens over time

A model in the house is not a device you set up and forget

Your own server is not a subscription that somebody else keeps current for you. Who runs it, and what happens if you want to move on later, belongs in the conversation now and not in the second year.

Keeping it current

A new model first runs alongside the old one and answers the same questions. Only when it visibly answers better do we switch over. Nobody on the outside forces the timing on you.

Who runs it

Either your own IT people, with documentation and training from us, or we do it as part of a maintenance arrangement. Both are common. The difference is who gets the alert at night.

If you want to move later

The documents, the search and the access rules are yours and sit in ordinary formats. Moving to another model, or later to a provider after all, is a relocation rather than a rebuild.

Your own server, or a provider’s API?

We build both. The question is not which route is better, but what your material allows.

Where the data ends up

On your own server

On your machine or in a Swiss data centre. Nothing goes outside at all.

Through a provider’s API

The question and the passages sent with it go to the provider’s servers.

Training on your content

On your own server

Does not arise, because there is nobody on the receiving end.

Through a provider’s API

Depends on the provider’s terms. Business plans usually exclude it, but it still has to be read.

Time to get started

On your own server

Longer: hardware, installation, access rules and a trial on your own documents.

Through a provider’s API

Short. One account is enough and the first trial runs the same week.

Outlay against usage

On your own server

A one-off set-up, then running costs and electricity, whatever the volume of questions.

Through a provider’s API

Nothing to buy, but an invoice per question that grows along with the use.

Computing power available

On your own server

As much as your machine has in it. More power means more hardware.

Through a provider’s API

The largest models are there immediately, with nothing for you to upgrade.

Updating the model

On your own server

You decide when to change, and you may stay on the version you have.

Through a provider’s API

The provider sets versions and retirements, and you follow along.

Fit with professional secrecy

On your own server

The material stays exactly where it already sits today.

Through a provider’s API

Requires a data processing agreement and, in case of doubt, the consent of the people concerned.

We build both routes and we earn no more on one than on the other. What fits is decided by your data and by how quickly you need it, not by our preference.

How a project like this runs

Four stages, and the first one is not about technology at all.

  1. 1

    Look at the data

    We go through with you which material is in play at all, and which of it is genuinely sensitive. It is often less than feared, and then the installation gets smaller.

  2. 2

    Choose the model and the set-up

    Two or three open models answer your real questions from your real documents. You read the answers side by side and decide on that basis, not on a specification sheet.

  3. 3

    Install it, with access rules and a log

    Installation in your house or in the Swiss data centre, tied into your existing accounts, with a log of every exchange and no route to the outside.

  4. 4

    Training and upkeep

    Your people learn what the assistant is good for and what it is not. After that your own IT runs it with our documentation, or we take on the ongoing upkeep.

What an installation like this is built with

Every piece of it works without a connection to the outside.

Open models

Llama, Mistral, Qwen, Gemma

Execution

Ollama, vLLM, llama.cpp

Document search

PostgreSQL with pgvector, Qdrant

Operation

Docker, your own server or a Swiss data centre

Frequently asked questions
What does AI in the house cost?

It depends on the hardware, on the size of the model and on how much material is to become searchable. That is why we do not name a figure before we know. At the free initial consultation we settle the scope, and you then receive a fixed price for exactly that scope. We work through both the bought machine and the one rented in a Swiss data centre, so you see the two options side by side.

Free initial consultation
Do we need IT people of our own for this?

No. Your own IT can run it if that is what you want, and the documentation and training come from us. If there is nobody in the house, we take on the upkeep: updates, monitoring and keeping the installation reachable. What stays with you either way is the data.

Is an open model worse than the big ones?

On general knowledge and on free writing, the largest models are ahead. On your questions about your own material something else counts: that the right passage is found and reproduced faithfully. There a mid-sized open model usually closes the gap. We test that on your own questions rather than simply asserting it.

Does the machine have to stand in our building?

No. A computer with a graphics card in a Swiss data centre is the usual middle way: nothing to buy, no server room, and the data still stays in the country and under your control. Whether that satisfies your internal rules is something we look at beforehand, because some directives insist on your own premises.

What happens if it fails?

The same question as for any server you already run: backups, a spare part and a plan for how long you could manage without it. We put that in writing before the installation goes live. If the machine sits in a data centre, that part is the operator’s job in any case.

Can we still move to a provider later?

Yes, and the other way round just as easily. The documents, the search and the access rules sit in such a way that only the part that writes the answer gets swapped. That is a changeover, not a second project. It is exactly why we build it that way from the start.

Tell us which material is not allowed to leave the building.

At the free, no-obligation initial consultation we look at your data and your internal rules and tell you honestly whether your own server is needed or whether something simpler will do. We reply within 24 hours.

Book an initial consultation

Just a quick question? The AI assistant in the bottom right answers straight away.

Contact

Ready for the summit?

Tell us about your project. We reply within 24 hours with an honest assessment. Free and without obligation.

hallo@aurphi.ch
AurPhi MoukrimMarktstrasse 18, 8853 Lachen SZ, SwitzerlandPostal address, no walk-in service. Appointments by arrangement.