THE MACHINE IN DETAIL

One appliance with three layers, one vendor.

The zenpAI appliance brings hardware, a local model and agent orchestration together into one system running on your own infrastructure. Security architecture, stack, hardware classes and operations in detail.

§01

Security as architecture.

NO ENDPOINT · NO FOREIGN JURISDICTION

Security sits in the architecture here: no shared cloud hardware, no external inference, no third-party administrators on your data. Your data stays with you.

GDPR DSGVO Meets the strict GDPR & EU AI Act requirements. No US CLOUD Act, no third country.
01

0% cloud exposure

All inference, all model weights, all training pipelines remain on your hardware. No sub-processor, no third country.

02

Local audit log

Every login, every tool call, and every relevant response is logged locally. This creates traceability for internal controls, compliance processes, and AI Act documentation.

§02

The stack.

FROM APPLICATION TO HARDWARE

Seven layers, one responsibility. At the bottom a hardened Linux with tuned drivers, above it Kubernetes on GitOps principles, versioned and auditable. The agent layer sits on top: it delegates tasks, keeps state, drives the flow and sets the governance while doing it. Proven datacenter software, not a hobby rig.

§02 · B · ARCHITECTURE

What actually runs in the rack, from the application down to the hardware.

APPLICATIONS User interfaces people work with
OpenWebUIStreamlitJetBrains JunieMCP-Clients
AGENT LAYER Orchestrates processes and keeps context
LangGraphModel Context ProtocolOAuth / LDAP / SSOlocal audit log
INFERENCE & MODEL LAYER Runs local AI models efficiently on the GPU
vLLMllama.cppTensorRT-LLMSGLangLlamaQwenDeepSeekGLMKimiMistralLoRA-Pipeline
PLATFORM Deployment, updates, and operation
KubernetesArgoCDHelmGitOpsHardened LinuxNVIDIA drivers / CUDAcontainerd
HARDWARE GPU server in your rack
NVIDIA Enterprise-GPUDELL PowerEdgeAMD EPYCredundant PSU
§03

Hardware classes.

ONE SIZING MODEL · VRAM · TOPS · USERS · kW

Three build sizes, matched to your load. The standard configuration builds on the current Blackwell generation. Final GPU, VRAM and PSU redundancy we set together, based on your load and latency profile.

SFOUNDATION / PILOT
GPU
1× NVIDIA RTX PRO 6000 · Blackwell
VRAM
96 GB
AI performance
~4,000 TOPS
Power max
~1 kW
Users
1–10
MPROFESSIONAL TEAM
GPU
4× NVIDIA RTX PRO 6000 · Blackwell
VRAM
384 GB
AI performance
~16,000 TOPS
Power max
~3 kW
Users
10–60
LCUSTOM ENTERPRISE
GPU
8× NVIDIA RTX PRO 6000 · Blackwell
VRAM
768 GB
AI performance
~32,000 TOPS
Power max
~5.5 kW
Users
60–200
IndividualMAXIMUM SCALE
GPU
Individually configured
VRAM
as required
AI performance
as required
Power max
individual
Users
200+

User counts depend on model and load. The appliance grows with you: we add GPU capacity when you need it. On request.

§04

How we deliver.

FROM FIRST CALL TO PRODUCTION · IN PHASES

Anyone can buy hardware. Integration, data remediation and agentic workflows are where the effort goes. We take that off your hands, from process assessment through to running operations, in clearly separated phases and billed per phase. How long it takes depends on the customer and the scope. If the first phase shows it is too early, we say so, before any hardware is ordered.

BEFORE KICK-OFF

Consulting & pilot

Assessment: where does it hurt, which data, which tools? We pick the first use case and show a concrete result fast. You can stop after that, no discussion.

PHASE 1

Analysis & data

Map data paths, define the goals. Most of it is data remediation: making scattered records production-ready.

PHASE 2

Build & fine-tuning

The appliance is racked and connected. The model is fine-tuned on your domain and quantized for your GPU.

PHASE 3

Orchestration & handover

Agent layer set onto your process, pilot operation, training for your IT, full handover. You take over, we stay on call. No lock-in.

Sounds like your project? Book an intro call

§05

Operations after go-live.

OPERATIONS · OPTIONAL

An on-premise AI has to be operated: monitored, updated, maintained. By default we hand everything over to you so your own IT can run it. If you prefer, we take operations on entirely, remotely, for a fixed monthly price. No token fees, no usage billing, no surprises.

INCLUDED IN THE SERVICE
  • Continuous 24/7 monitoring: GPU, memory, latencies, system status
  • Security and system updates for the hardened Linux stack
  • Current, curated models: we evaluate new generations continuously and roll them out
  • Remote support and error analysis
  • Hardware early warning and upgrade recommendations
ADD-ONS · PROJECT-BASED
  • Re-training and RAG optimization on your own data
  • Custom development and integration
  • On-site visits by arrangement

Not part of the monthly price: the hardware as a one-time purchase (see §04) and, for larger configurations, cooling. We quote the monthly price on request. It depends on hardware class and the response time you need.

One appliance, one vendor. Let us talk.

We set the hardware class, model family and integration depth together with you, once we understand your data, your systems and your first use case. There is no configurator for that here.

Book intro call →
Book an intro call →