The zenpAI appliance brings hardware, a local model and agent orchestration together into one system running on your own infrastructure. Security architecture, stack, hardware classes and operations in detail.
Security sits in the architecture here: no shared cloud hardware, no external inference, no third-party administrators on your data. Your data stays with you.
All inference, all model weights, all training pipelines remain on your hardware. No sub-processor, no third country.
Every login, every tool call, and every relevant response is logged locally. This creates traceability for internal controls, compliance processes, and AI Act documentation.
Seven layers, one responsibility. At the bottom a hardened Linux with tuned drivers, above it Kubernetes on GitOps principles, versioned and auditable. The agent layer sits on top: it delegates tasks, keeps state, drives the flow and sets the governance while doing it. Proven datacenter software, not a hobby rig.
What actually runs in the rack, from the application down to the hardware.
Three build sizes, matched to your load. The standard configuration builds on the current Blackwell generation. Final GPU, VRAM and PSU redundancy we set together, based on your load and latency profile.
User counts depend on model and load. The appliance grows with you: we add GPU capacity when you need it. On request.
Anyone can buy hardware. Integration, data remediation and agentic workflows are where the effort goes. We take that off your hands, from process assessment through to running operations, in clearly separated phases and billed per phase. How long it takes depends on the customer and the scope. If the first phase shows it is too early, we say so, before any hardware is ordered.
Assessment: where does it hurt, which data, which tools? We pick the first use case and show a concrete result fast. You can stop after that, no discussion.
Map data paths, define the goals. Most of it is data remediation: making scattered records production-ready.
The appliance is racked and connected. The model is fine-tuned on your domain and quantized for your GPU.
Agent layer set onto your process, pilot operation, training for your IT, full handover. You take over, we stay on call. No lock-in.
Sounds like your project? Book an intro call →
An on-premise AI has to be operated: monitored, updated, maintained. By default we hand everything over to you so your own IT can run it. If you prefer, we take operations on entirely, remotely, for a fixed monthly price. No token fees, no usage billing, no surprises.
Not part of the monthly price: the hardware as a one-time purchase (see §04) and, for larger configurations, cooling. We quote the monthly price on request. It depends on hardware class and the response time you need.
We set the hardware class, model family and integration depth together with you, once we understand your data, your systems and your first use case. There is no configurator for that here.