Business

What we do AI InfrastructureLLM & Token PlatformApplied AI

Company

Technology & Operations Case Studies News Company Careers Contact

Case studies

Case studies

Published under confidentiality agreements, with client names and identifying details removed. Figures are measured results from each engagement.

Manufacturing

Making forty years of design drawings searchable by the engineers themselves

Manufacturer, Client A (approx. 1,000 employees)

−64%

Engineering hours spent searching drawings

Challenge

Every design change required finding comparable historical drawings, but naming conventions differed by era and annotations included handwritten scans. Asking senior engineers verbally had become the de facto search mechanism, concentrating knowledge in a few people.

Approach

We built a multimodal retrieval platform over the drawing archive. Working with the engineering team, we established a normalisation dictionary for drawing-number and part-name variation, and implemented hybrid retrieval combining OCR output from annotations with drawing metadata. Every result links back to the source drawing, so no AI output is trusted without grounding.

Outcome

Hours spent identifying comparable drawings fell 64% on average. Verbal queries to senior engineers dropped from roughly 210 to 58 per month, removing the dependency on individuals.

Key technologies
On-premise GPU (8× L40S), Japanese embedding model, hybrid retrieval (BM25 + vector), permission-aware filtering
Duration
4-week assessment plus 5-month build
Financial services

Consolidating scattered departmental AI contracts behind one gateway

Financial services provider, Client B

−52%

Token spend

Challenge

Three years of independently signed departmental AI contracts had left the organisation with no view of either total spend or actual usage. Internal audit had asked for unified logging, which the fragmented contracts made impossible.

Approach

We built an LLM gateway for one group company inside the client's VPC, adopting an OpenAI-compatible schema so existing applications could migrate by changing only the endpoint. Requests are routed automatically by use case — routine enquiries to smaller models, difficult analysis to frontier models — with prompt caching applied throughout.

Outcome

Monthly token spend fell 52%. Audit logging for all requests is now unified and satisfies the audit requirement. Adoption grew from three departments to eleven, and departmental chargeback made cost allocation a discussable subject.

Key technologies
LLM gateway (in-house architecture), automatic model routing, prompt caching, departmental chargeback, audit logging platform
Duration
3-week assessment plus 4-month build, ongoing operations
Research

A 128-GPU training cluster the researchers do not have to operate

Research institution, Client C

−71%

Training job queue time

Challenge

Researchers were doubling as cluster administrators, and research halted every time an incident occurred. Job priority control effectively did not exist, so large jobs blocked small ones for long periods.

Approach

Delivered jointly with a data centre operator, the 128-GPU cluster runs on an InfiniBand NDR fabric. YOLAIZ designed priority control and preemption in Slurm, fair-share allocation by research group, and automatic isolation and rescheduling on node failure, and holds first response in operations. The entire configuration was delivered as Terraform and Ansible code, ready for transfer to the institution.

Outcome

Average training job queue time fell 71%, and researcher time spent on cluster operations went to effectively zero. Nine node failures occurred in the first twelve months of operation; all were resolved by job resubmission alone, with no impact reaching researchers.

Key technologies
128× NVIDIA H100, InfiniBand NDR 400G, Slurm with fair-share, Lustre, Prometheus/Grafana, Terraform/Ansible
Duration
3-month design plus 4-month build, ongoing operations
SaaS

p95 latency from 2.4s to 0.9s, with fewer GPUs

SaaS provider, Client D

−35%

GPUs required

Challenge

The LLM feature embedded in the client's product was slow enough to start appearing in churn reasons. Adding capacity had held the line, but it was eroding gross margin and further expansion would not clear a management decision.

Approach

Profiling the existing deployment identified batching strategy and KV cache handling as the bottleneck. We rebuilt the configuration around vLLM continuous batching and paged attention, combining FP8 quantisation with speculative decoding. The application's prompt structure was redesigned in parallel so that the shared prefix could be cached effectively.

Outcome

p95 latency improved from 2.4 seconds to 0.9 seconds while GPU count fell 35%, improving gross margin. Response speed no longer appears in churn reasons.

Key technologies
vLLM, FP8 quantisation, speculative decoding, prompt caching, NVIDIA H100
Duration
2-week assessment plus 7-week optimisation
Public sector

Supporting enquiry handling with no data leaving the building

Public service provider, Client E

−48%

Time to first response

Challenge

Handling residents' enquiries required consulting a wide range of regulations and notices, and differences in staff experience translated directly into differences in response time. Sending data to external APIs was not an available option under the organisation's rules.

Approach

Working as subcontractor to a prime integrator, we built a fully on-premise inference environment a RAG platform over regulations, notices and historical responses. Draft answers always carry links to the governing provisions, and the workflow requires a staff member to confirm before anything is issued. The system enforces that AI never finalises a response on its own.

Outcome

Time to first response fell 48% on average, and variance in response time by staff experience narrowed. Citing the governing provisions also made responses easier to verify.

Key technologies
On-premise GPU, open-weight model, RAG platform, mandatory source citation, audit logging
Duration
4-week assessment plus 6-month build
Manufacturing

No longer buying GPUs that exist only for the peak

Precision equipment manufacturer, Client F

−29%

Annual TCO of compute

Challenge

Compute demand concentrated during new product design periods, and GPUs were purchased to match those peaks. Average utilisation across the year sat at 31%, leaving the return on investment unexplainable.

Approach

Analysing twenty-four months of demand, we designed a hybrid topology with baseline load on the client's own capacity and overflow bursting to public cloud. Routing rules were defined by job characteristics, with placement constraints ensuring sensitive workloads never leave the client's own cluster.

Outcome

Annual total cost of ownership for compute fell 29%, while average utilisation of the client's own cluster rose from 31% to 78%. Queue times during design periods did not increase.

Key technologies
On-premise GPU cluster, public cloud burst, Kubernetes + Kueue, sensitivity-aware placement constraints
Duration
4-week assessment plus 3-month build

Published with each client's consent, edited so that the organisation cannot be identified. Detailed conditions can be discussed individually.

Bring us in while it is still an idea.

The most common questions we receive are "how many GPUs should we buy" and "should we own or rent". We welcome conversations long before requirements are settled.

Go to the contact form