Run production AI inside your own cloud.
Gerniq is a Kubernetes-native serving platform for document, vision and language models. The same code runs on a laptop, in your VPC or in your data centre, and your data never leaves your environment.
Private preview · Runs on any Kubernetes: EKS, AKS, GKE, OKE or on-prem
$ make up # local kind cluster + sample models ✓ Gerniq control plane ready ✓ Gateway listening on https://localhost $ gerniq model push ./invoice-extractor --runtime onnx $ gerniq deploy invoice-extractor --profile cpu-small ✓ invoice-extractor ready (1/1 replicas) $ curl -s https://localhost/v1/pipelines/invoice/run \ -H "Authorization: Bearer $TOKEN" -F file=@invoice.pdf { "invoice_number": "INV-10423", "invoice_date": "2026-09-14", "total": "48250.00", "currency": "INR" }
Illustrative session from the private preview.
Enterprise AI stalls in three places.
“Our data can't leave.”
Customer documents, health records and financial files can't be sent to an external API, so promising AI pilots never reach production.
“Every team deploys models differently.”
Each model ships with its own server, scripts and endpoints. Nobody can say what is running, where, or at what cost.
“GPU spend is unpredictable.”
LLMs get GPUs by default, even when a CPU model would do. Idle capacity and cold starts eat the budget.
From model file to production API in three steps.
One registry, one deployment model and one API contract for every model your teams ship.
Register
Push a model and its runtime profile to the Gerniq registry. Artifacts stay in your own S3-compatible storage.
gerniq model push ./kyc-ocr --runtime triton
Deploy
Declare what should run. The controller rolls it out, scales it and reports health back to the registry.
gerniq deploy kyc-ocr --profile gpu-a10 --min 0
Call
Use one API: OpenAI-compatible for LLMs and vision-language models, KServe v2 for classic models, and pipelines for multi-step work.
POST /v1/chat/completions POST /v2/models/kyc-ocr/infer
Everything between your model and your users.
Gerniq adds the production layer that open-source runtimes leave out, and keeps it on open protocols.
Model registry
Versions, formats, owners and lifecycle for every model, backed by PostgreSQL and your object storage.
Inference gateway
Authentication, tenant quotas, rate limits, request validation and streaming responses in one place.
Pipelines
Chain models into versioned workflows, such as layout, then OCR, then field extraction, with partial-failure handling.
Async jobs
Submit multi-page documents as jobs. Workers fan out pages, retry safely and write results to your storage.
Autoscaling
Scale on latency, concurrency or queue depth. GPUs for LLMs, CPUs for classic models, scale-to-zero when idle.
Observability
OpenTelemetry traces, Prometheus metrics and per-model dashboards for throughput, p95 latency, errors and saturation.
Bring the models you already use.
Gerniq runs standard open-source runtimes, so your models are never locked into ours.
What you can serve
- OCR and layout detection
- Key-value and table extraction
- Classification and object detection
- Vision-language models
- Open-weight LLMs
- Your own fine-tuned and custom-trained models
Same code from laptop to production.
Gerniq speaks only open protocols: Kubernetes, S3, Kafka, PostgreSQL, OIDC and OpenTelemetry. Moving between environments changes configuration, not code.
Built first for document-heavy work.
Banks, insurers and hospitals process millions of pages a year. Gerniq runs the whole pipeline next to the data.
KYC and onboarding
Read identity documents, address proofs and forms, extract fields and flag mismatches for review.
Loan and credit files
Classify and extract bank statements, salary slips and tax returns across hundreds of pages.
Insurance claims
Pull policy numbers, amounts and diagnoses from claim forms, hospital bills and discharge summaries.
Medical records
Digitise lab reports and prescriptions without data leaving the hospital network.
Your data stays inside your walls.
Gerniq is software you run, not a service you send data to. Models, requests and results stay in your cluster and your storage.
- Runs in your VPC, data centre or air-gapped network
- Single sign-on through your identity provider (OIDC)
- Role-based access and per-team quotas
- Audit logs for model changes and API access
- TLS everywhere, network policies and signed images
Now onboarding three design partners.
We are working with a small number of teams in banking, insurance and healthcare to take one document workflow to production in 6 to 8 weeks, inside their own environment.
- Hands-on engineering from the founding team
- Your workflow live in your cluster, against success criteria agreed up front
- Pilot fee credited against your first-year licence
- Direct influence on the roadmap
Start free. Prove it on one workflow. Then scale.
Community
Evaluate the full platform on a laptop or a dev cluster.
Pilot
One workflow live in your environment in 6 to 8 weeks.
Enterprise
Production use with SSO, audit, SLA and a named engineer.
Questions platform teams ask first
Does any of our data leave our network?
No. Gerniq runs entirely inside your Kubernetes cluster and uses your own storage, database and identity provider. Usage telemetry is optional and off by default.
Do we need GPUs?
Not for everything. Classic document and vision models run well on CPU. GPUs are needed only for vision-language models and LLMs, and Gerniq can scale those to zero when idle.
Which models are supported?
Any model that runs on ONNX Runtime, NVIDIA Triton, vLLM or a Python server, including open-weight LLMs and your own custom-trained models.
How is this different from running KServe or vLLM ourselves?
Gerniq builds on those open-source projects and adds what a production platform needs: a model registry, an inference gateway with auth and quotas, pipelines, async jobs, metering and dashboards. You start from a working platform instead of assembling one over months.
Which Kubernetes distributions do you support?
Gerniq targets conformant Kubernetes, including EKS, AKS, GKE, OKE and on-prem distributions. We confirm your exact versions during the pilot scoping call.
What support do we get?
Design partners work directly with the founding team. Enterprise licences include an SLA and a named engineer.
See Gerniq running on your documents.
Book a 30-minute call. We will scope a pilot around one workflow and your environment.