Business

What we do AI InfrastructureLLM & Token PlatformApplied AI

Company

Technology & Operations Case Studies News Company Careers Contact

News

News

Press releases, technical notes and corporate announcements.

2026-07-28

Corporate

GPUs under our operation pass 400

Since our founding in May 2024, the total number of GPUs entrusted to us for operation has passed 400. We continue to expand our arrangements with partner data centres.

2026-06-17

Technical

Measured results from speculative decoding in production inference

Applying draft-model speculative decoding to 70B-class models, we measured the effect on output token cost and p95 latency, observing a 31–44% p95 improvement depending on conditions.

2026-05-20

Press release

Departmental chargeback added to the LLM gateway

Requests are now tagged by department, use case and model, with monthly cost allocation reports generated automatically. Available to existing clients at no additional charge.

2026-04-08

Technical

Validation results for direct liquid cooling in high-density racks

With the cooperation of a partner data centre, we validated PUE and GPU temperature distribution with direct liquid cooling in configurations above 60 kW per rack. Comparative data against air-cooled configurations is published as a technical note.

2026-03-02

Corporate

Applied AI team formed

To expand our RAG platform and operational AI agent capacity, we have formed a dedicated Applied AI team with assigned solutions architects.

2026-01-22

Press release

Private inference environment for a financial services provider enters full operation

The private inference environment built in December 2025 has entered full operation as the platform for all departments of a group company, expanding from three departments at launch to eleven.

2025-11-11

Technical

GPU allocation design in mixed Slurm and Kubernetes environments

A technical note on how we design priority control and preemption where training jobs and inference workloads share a single cluster.

2025-08-05

Press release

Inference optimisation service launched

We now offer contracted improvement of throughput, latency and token cost on vLLM and TensorRT-LLM. Assessment-only engagements against existing deployments are also available.

Bring us in while it is still an idea.

The most common questions we receive are "how many GPUs should we buy" and "should we own or rent". We welcome conversations long before requirements are settled.

Go to the contact form