One serving platform for every model you run.
Gerniq gives platform teams a registry, a gateway, pipelines and autoscaling on top of proven open-source runtimes, deployed in your own Kubernetes cluster.
Every component runs inside your cluster.
Stateful services can run in-cluster for development or use your managed equivalents in production. Only configuration changes between the two.
Solid boxes are Gerniq components or open-source runtimes it manages. The bottom band lists services Gerniq connects to through open protocols.
What each part does, and what it is built on.
| Component | What it does | Built on |
|---|---|---|
| Model registry | Stores models, versions, artifact locations, formats, owners and lifecycle state (draft, active, deprecated). | PostgreSQL, S3 API |
| Deployment controller | Reconciles declared deployments into running, autoscaled model servers and writes readiness back to the registry. | Kubernetes custom resources, KServe |
| Runtime catalog | Version-pinned runtimes and resource profiles such as cpu-small, cpu-large and gpu-a10. | ONNX Runtime, NVIDIA Triton, vLLM |
| Inference gateway | Authenticates callers, resolves tenants, enforces quotas and rate limits, validates requests, routes to models and streams LLM output. | OIDC / JWT, Gateway API |
| Pipeline orchestrator | Runs versioned multi-model workflows with per-page fan-out, bounded concurrency and partial-failure handling. | Gerniq |
| Jobs API and workers | Accepts large or multi-page inputs as jobs, splits them, retries safely and writes results to your storage. | Kafka protocol, S3 API |
| Autoscaling | Scales on latency, concurrency and queue lag, with optional scale-to-zero for GPU models. | KEDA, cluster autoscaler |
| Observability | Per-model dashboards and alerts for throughput, p50/p95 latency, errors and GPU saturation. | OpenTelemetry, Prometheus, Grafana |
| Security | TLS certificates, network policies, pod security standards and secrets from your vault. | cert-manager, External Secrets Operator |
One contract per kind of model.
Your applications call the gateway, never a model pod directly. Endpoints come from the registry, not from naming conventions.
- OpenAI-compatible for LLMs and vision-language models, including streaming
- KServe v2 (Open Inference Protocol) for classic models
- Pipelines and Jobs for multi-step and multi-page work
- CLI and Python SDK for registering and deploying models
# LLMs and vision-language models POST /v1/chat/completions { "model": "doc-vlm", "messages": [ ... ] } # Classic models (Open Inference Protocol) POST /v2/models/invoice-kv/infer # Multi-page documents as async jobs POST /v1/jobs { "pipeline": "claims", "input": "s3://claims-inbox/2026-09/" } GET /v1/jobs/{id}
Laptop, cloud or data centre: same platform.
| Environment | How Gerniq runs there |
|---|---|
| Laptop (kind) | CPU-only cluster with in-cluster object storage, PostgreSQL, Kafka-compatible queue and identity provider. LLMs use a small model or a mock. |
| Managed Kubernetes EKS, AKS, GKE, OKE | Uses your managed object storage, database, Kafka service and identity provider. Separate CPU and GPU node pools, cluster autoscaling and GitOps delivery. |
| On-prem Kubernetes | Runs on your own servers and GPUs with your existing storage, database and directory. |
| Air-gapped | Offline bundle of container images, Helm charts and model weights, loaded into your internal registry. Available to design partners on request. |
Run the whole platform on your laptop.
One command brings up a local kind cluster with ingress, storage, database, queue, identity, the gateway and sample models.
The quickstart is shared with design partners and early-access users during the private preview.
Request quickstart access$ make up ✓ kind cluster (1 control plane, 2 workers) ✓ ingress, object storage, PostgreSQL, queue, identity ✓ gateway, registry, controller ✓ sample models registered and deployed $ gerniq model list $ curl -k https://localhost/v2/health/ready
Illustrative output. Exact steps ship with the quickstart.
Want to see it on your cluster?
We will run a scoping call, then deploy Gerniq in your staging environment as part of a pilot.