Kalufs
Dedicated AI. Customer control.

Your work.
Your models.
Your terms.

Cost-effective physical hardware dedicated to your organisation alone. Never shared with another customer.

Discuss your needs

We bring together dedicated hardware, carefully selected models and the expertise to make them work well. You decide where it runs, how your data is handled and which capabilities your team needs.

  • Analyse and process confidential or proprietary data
  • Deploy sovereign agentic workflows
  • Generate and edit media showcasing classified or unreleased products
Control that means something

Keep the decisions that matter.

Agree the operating arrangements that matter to your team: data policies, model capabilities and controlled changes.

Your data and policies

You own your data and decide audit logging, retention and usage policies. Your data never leave your dedicated hosts, so Kalufs does not inspect your data unless you let us.
Why European soil is no guarantee against foreign interests acquiring your data →
Why privacy is not the whole confidentiality question →

Models that can do your work

Legitimate work can involve sensitive material. We help you select and validate models, including reduced-refusal or abliterated variants, when your workflows need them. You set the usage policies.
When safeguards block legitimate work →

Changes made deliberately

We agree model changes with you and validate them against your workloads. Kalufs configures and tunes the agreed models to make effective use of the hardware.
Why model lifecycle is part of the service →

One gateway. Your configuration.

Different capabilities.
A shared way in.

Your applications request an agreed model or role alias. We configure what stays ready and what loads on demand with you.

Run models from families such as DeepSeek, Gemma, GLM, Ideogram, Kimi, Krea, LTX, MiMo, Minimax, Mistral, Muse, Qwen and YuE on your own terms.

Example A

An engineering team

One dedicated node · eight GCDs · 512 GB VRAM
  1. GCD 0Reasoning model · shard
  2. GCD 1Reasoning model · shard
  3. GCD 2Reasoning model · shard
  4. GCD 3Reasoning model · shard
  5. GCD 4Prose model · shard
  6. GCD 5Prose model · shard
  7. GCD 6
    Image model XImage model Y
  8. GCD 7
    3D asset generation modelOCR model
1 TB system RAM

Model loading, host workloads and optional KV-cache offload share this capacity.

Reasoning model, prose model, 3D asset generation and OCR remain resident in this example. Image model X and Image model Y are exposed as separate model endpoints; the selected image model is swapped into GCD 6 on demand.

Example B

A research and creative studio

One dedicated node · eight GCDs · 512 GB VRAM
  1. GCD 0Prose model · shard
  2. GCD 1Prose model · shard
  3. GCD 2Prose model · shard
  4. GCD 3Prose model · shard
  5. GCD 4Video generation model · shard
  6. GCD 5Music generation model
  7. GCD 6
    Image model XImage model Y
  8. GCD 7
    Text-to-speech model3D asset generation model
1 TB system RAM

Model loading, host workloads and optional KV-cache offload share this capacity.

Prose model, video generation model, music generation model, text-to-speech model and 3D asset generation model remain resident in this example. Image model X and Image model Y are exposed as separate model endpoints; the selected image model is swapped into GCD 6 on demand.

Example C

A legal team

One dedicated node · eight GCDs · 512 GB VRAM
  1. GCD 0Prose model · shard
  2. GCD 1Prose model · shard
  3. GCD 2Prose model · shard
  4. GCD 3Prose model · shard
  5. GCD 4
    Image analysis modelVideo analysis model
  6. GCD 5Custom legal research + citations model
  7. GCD 6
    OCR modelSpeech transcription model
  8. GCD 7
    Translation modelStructured extraction model
1 TB system RAM

Model loading, host workloads and optional KV-cache offload share this capacity.

Prose model, custom legal research + citations model, OCR model, speech transcription model, translation model and structured extraction model remain resident in this example. Image analysis model and Video analysis model are exposed as separate model endpoints; the selected analysis model is swapped into GCD 4 on demand.

Each tile is one GPU compute die (GCD), not a whole graphics card or a virtual security boundary. Repeated colours show one model distributed across multiple GCDs. Occupied tiles identify allocations. Model names and labels are illustrative.

Diffusion models can share a GCD when their combined weights and execution memory fit. Typical diffusion generation does not need the growing autoregressive KV cache used by an LLM, but activations, attention buffers and image/video decoding still need memory.

Supported serving stacks can offload LLM KV cache to system RAM. RAM remains on the same dedicated host.

Exact models, memory use and loading arrangements are agreed for your work.

A stable name. An agreed upgrade.

Your application: “reasoning”Current agreed modelValidated replacement

As new, better models are released or your needs change, we can replace models at your request. We ensure they run efficiently on your dedicated hardware and can help customise, fine-tune and adapt foundation models to your needs. In every case, we provide a stable, agreed upgrade path.

Room to grow

Start with one node.
Expand with your work.

Each node has ConnectX-6 200GbE connectivity. Add dedicated nodes as demand grows: to serve more work in parallel, or to accommodate a larger model across nodes with a compatible serving stack.

More work at once

Run model replicas or different workloads on additional nodes.

Node A · model replica
Node B · additional replica
ConnectX-6 · 200GbE network connectivity

Distribute requests across agreed deployments. Additional capacity can support more concurrent users and jobs.

A larger model across nodes

Distribute one larger model across multiple nodes.

Node A · model shards
Node B · remaining model shards
ConnectX-6 · 200GbE network connectivity

Use a supported distributed serving configuration when one node is not enough.

We agree and validate node count, model partitioning, network topology and performance for your workload. Additional nodes follow the same agreed operating arrangements.

Placement is your choice

Close to your work.
Inside your boundary.

Choose your deployment

Choose an on-premises deployment or hosting in Östersund (Sweden). Hosting in Imperia (Italy) is coming soon! Location is a separate choice from which models stay loaded.

On your premises

Dedicated hardware is physically installed at your premises. You can choose an air-gapped deployment for extra isolation.

Updates and modifications require physical access to the host. This can lengthen support response and resolution times. Maintenance arrangements and support expectations are agreed as part of the deployment.

A dedicated server rack inside a stylised office room.

Östersund

In the middle of Sweden, not far from Trondheim in Norway and near the mountains, Östersund is “Vinterstaden”—“the Winter City”. Our partner’s data centre is supplied with fossil-free hydropower.

A snow-covered city beside water, framed by mountain silhouettes.

Imperia

Coming soon! Our partner’s data centre in Imperia is under construction and will soon be operational.

On the Ligurian coast at the foot of the Ligurian Alps, less than an hour from the Nice metropolitan area in France, Imperia celebrates its slogan “il miglior clima d’Italia”—“Italy’s best climate”. Free cooling reduces cooling energy demand, while abundant sunshine supports fossil-free solar generation.

A sunlit Ligurian coastal town between mountains and sea.
From deployment to delivery

More than a host.
A reliable partner.

We help you put the hardware to work, from deployment and integration to ongoing support.

Kalufs offers consultancy in software development, agentic workflows, systems integration, data labelling, training operations, MLOps, red teaming, penetration testing, security monitoring, incident response and managed fine-tuning.

Start with the work you need to do.

A useful first conversation covers:

  • Your area of business, your data-sovereignty challenges and opportunities.
  • Your workflows and required model capabilities.
  • Where processing should happen and who may access the data.
  • The integrations, evaluations and support your team needs.

Interested? Send us an e-mail at info@kalufs.se for an informal conversation about your needs.

Read before you decide

Different reasons
to take control.

Articles on enduring questions, updated in place as the topics develop.

Your privacy

We value your privacy. We use no analytics cookies and do not track you across websites. We use Plausible Analytics only to understand aggregate website traffic, such as visitor numbers and general geographic locations. We respect Do Not Track and Global Privacy Control.