Business

What we do AI InfrastructureLLM & Token PlatformApplied AI

Company

Technology & Operations Case Studies News Company Careers Contact

Technology & Operations

Technology and operations

We publish the technologies we use, how engagements run, and the operational commitments that apply in production — so that our actual scope can be judged early in your evaluation.

Technology stack

Technology stack

Limited to technologies with production operating history in our engagements.

Compute & hardwareNVIDIA H100 / H200 / B200 / L40S / A100, InfiniBand NDR and HDR, RoCEv2, liquid and high-density air-cooled racks
OrchestrationSlurm, Kubernetes, Kueue, Ray, Karpenter, Argo Workflows
StorageLustre, WEKA, Ceph, MinIO, NVMe-oF
Inference & trainingvLLM, TensorRT-LLM, SGLang, DeepSpeed, FSDP, Megatron-LM, PyTorch
Model operationsMLflow, Weights & Biases, LangFuse, in-house evaluation harness
Platform & IaCTerraform, Ansible, Pulumi, GitHub Actions, ArgoCD
ObservabilityPrometheus, Grafana, Loki, Tempo, DCGM Exporter, OpenTelemetry
CloudAWS, Google Cloud, Microsoft Azure, Sakura Internet, domestic data centre operators

Engagement process

Engagement process

From first conversation to production, engagements normally follow this sequence. Each stage is contracted so that it can be stopped at its boundary.

01

First conversation

No charge, one or two sessions

We ask about the current situation, existing assets and rough budget. Requirements need not be settled. If we are not the right firm for the work, we say so.

02

Assessment

2–4 weeks

Workload measurement, review of existing infrastructure, cost modelling, and several architecture options. You are free to take these deliverables to another vendor.

03

PoC & design

6–10 weeks

Performance and unit cost measured on a reduced configuration and fed into the production design. The numeric criteria for proceeding are agreed at this stage.

04

Production build

3–6 months

Procurement, installation, infrastructure-as-code, monitoring and acceptance testing. The entire configuration is delivered as code.

05

Operations & improvement

Monthly

Monitoring and first response, scheduled uptime and cost reviews, continuous architectural revision. Where in-housing is desired, a transfer plan runs in parallel.

Operations and service levels

Operations and service levels

Standard levels; individual contracts may vary.

Monitoring24/365 automated monitoring of metrics, logs and GPU health
Support hoursWeekdays 09:30–18:00 JST; out-of-hours cover arranged case by case with partners
First responseWithin 30 minutes for critical incidents, one business hour otherwise
Uptime target99.5% monthly, excluding planned maintenance
ReportingMonthly report covering uptime, job statistics, cost and improvement proposals
Change managementEvery configuration change recorded and reviewed as a pull request

Security posture

Security posture

AI platforms touch confidential information through two channels: training data and inference input. Our baseline is as follows.

01

Keeping data inside by default

For both inference and training, our first proposal is always an architecture that completes inside your control boundary. Where an external API is used, we state exactly what data leaves.

02

Permissions carried through to results

In RAG platforms, source document permissions are reflected in retrieval. Documents a user cannot access never become grounding for an answer — this is a design requirement, not a setting.

03

Every inference auditable

Who sent what to which model, and what came back. Recorded centrally at the gateway layer, with retention periods fixed by contract.

04

No secrets in deliverables

Infrastructure-as-code carries no credentials; references go through a secrets manager. Automated scanning runs before delivery.

The full Information Security Policy is available here.

On our size

One thing about scale, stated up front.

We are a small firm, and the number of engagements we can hold at once is genuinely limited. That is commercially inconvenient to admit, so we admit it here rather than later.

In exchange, the people you meet first are the people who stay through operations. Every architectural decision passes a review that includes the CEO. Nothing is lost to a handover.

Builds beyond 128 GPUs, and engagements requiring staffed 24/365 cover, are delivered jointly with data centre operators and systems integrators. Our part is the design, the operational design, and technical accountability inside that arrangement. We say which it is at the outset.

Bring us in while it is still an idea.

The most common questions we receive are "how many GPUs should we buy" and "should we own or rent". We welcome conversations long before requirements are settled.

Go to the contact form