01
LLM gateway
Commercial and open-weight models consolidated behind a single API endpoint, with cross-model failover, per-department rate limiting and unified audit logging as standard.
INFERENCE PLATFORM
Models keep multiplying, and both price and capability turn over within months. Wiring a business system directly to one model is a commitment to rewriting it every time. We place one replaceable layer in between.
01
Commercial and open-weight models consolidated behind a single API endpoint, with cross-model failover, per-department rate limiting and unified audit logging as standard.
02
KV cache strategy, continuous batching, quantisation (FP8 / INT4) and speculative decoding across vLLM / TensorRT-LLM / SGLang, tuned to your throughput and latency targets.
03
Every request tagged by model, department and use case, with monthly chargeback reporting generated automatically. Prompt caching and model routing lower unit cost without a perceptible quality drop.
04
LoRA / QLoRA through full-parameter tuning and continued pre-training, delivered end to end: data preparation, training, evaluation and deployment.
05
Japanese-language benchmarks plus evaluation sets built from your own operational data, so regressions on model updates surface as numbers rather than complaints.
06
Architectures that keep inference inside your control boundary, with prompts and outputs never leaving it. Available region-locked to Japan or fully on-premise.
Problems this solves
Platform specifications
| Models | 32 commercial and open-weight models, expanded continuously |
|---|---|
| Inference engines | vLLM, TensorRT-LLM, SGLang, Ollama (validation use) |
| Quantisation | FP8, INT8, INT4 (AWQ / GPTQ), KV cache quantisation |
| Gateway features | Model routing, failover, rate limiting, prompt caching, audit logging, departmental chargeback |
| API compatibility | OpenAI-compatible schema — existing clients migrate by changing the endpoint only |
| Deployment | YOLAIZ-managed, inside your VPC, or on-premise |
| First response | Within one business hour on business days; out-of-hours cover arranged case by case with partners |
The most common questions we receive are "how many GPUs should we buy" and "should we own or rent". We welcome conversations long before requirements are settled.