01
Keeping data inside by default
For both inference and training, our first proposal is always an architecture that completes inside your control boundary. Where an external API is used, we state exactly what data leaves.
Technology & Operations
We publish the technologies we use, how engagements run, and the operational commitments that apply in production — so that our actual scope can be judged early in your evaluation.
Technology stack
Limited to technologies with production operating history in our engagements.
| Compute & hardware | NVIDIA H100 / H200 / B200 / L40S / A100, InfiniBand NDR and HDR, RoCEv2, liquid and high-density air-cooled racks |
|---|---|
| Orchestration | Slurm, Kubernetes, Kueue, Ray, Karpenter, Argo Workflows |
| Storage | Lustre, WEKA, Ceph, MinIO, NVMe-oF |
| Inference & training | vLLM, TensorRT-LLM, SGLang, DeepSpeed, FSDP, Megatron-LM, PyTorch |
| Model operations | MLflow, Weights & Biases, LangFuse, in-house evaluation harness |
| Platform & IaC | Terraform, Ansible, Pulumi, GitHub Actions, ArgoCD |
| Observability | Prometheus, Grafana, Loki, Tempo, DCGM Exporter, OpenTelemetry |
| Cloud | AWS, Google Cloud, Microsoft Azure, Sakura Internet, domestic data centre operators |
Engagement process
From first conversation to production, engagements normally follow this sequence. Each stage is contracted so that it can be stopped at its boundary.
01
No charge, one or two sessions
We ask about the current situation, existing assets and rough budget. Requirements need not be settled. If we are not the right firm for the work, we say so.
02
2–4 weeks
Workload measurement, review of existing infrastructure, cost modelling, and several architecture options. You are free to take these deliverables to another vendor.
03
6–10 weeks
Performance and unit cost measured on a reduced configuration and fed into the production design. The numeric criteria for proceeding are agreed at this stage.
04
3–6 months
Procurement, installation, infrastructure-as-code, monitoring and acceptance testing. The entire configuration is delivered as code.
05
Monthly
Monitoring and first response, scheduled uptime and cost reviews, continuous architectural revision. Where in-housing is desired, a transfer plan runs in parallel.
Operations and service levels
Standard levels; individual contracts may vary.
| Monitoring | 24/365 automated monitoring of metrics, logs and GPU health |
|---|---|
| Support hours | Weekdays 09:30–18:00 JST; out-of-hours cover arranged case by case with partners |
| First response | Within 30 minutes for critical incidents, one business hour otherwise |
| Uptime target | 99.5% monthly, excluding planned maintenance |
| Reporting | Monthly report covering uptime, job statistics, cost and improvement proposals |
| Change management | Every configuration change recorded and reviewed as a pull request |
Security posture
AI platforms touch confidential information through two channels: training data and inference input. Our baseline is as follows.
01
For both inference and training, our first proposal is always an architecture that completes inside your control boundary. Where an external API is used, we state exactly what data leaves.
02
In RAG platforms, source document permissions are reflected in retrieval. Documents a user cannot access never become grounding for an answer — this is a design requirement, not a setting.
03
Who sent what to which model, and what came back. Recorded centrally at the gateway layer, with retention periods fixed by contract.
04
Infrastructure-as-code carries no credentials; references go through a secrets manager. Automated scanning runs before delivery.
On our size
We are a small firm, and the number of engagements we can hold at once is genuinely limited. That is commercially inconvenient to admit, so we admit it here rather than later.
In exchange, the people you meet first are the people who stay through operations. Every architectural decision passes a review that includes the CEO. Nothing is lost to a handover.
Builds beyond 128 GPUs, and engagements requiring staffed 24/365 cover, are delivered jointly with data centre operators and systems integrators. Our part is the design, the operational design, and technical accountability inside that arrangement. We say which it is at the outset.
The most common questions we receive are "how many GPUs should we buy" and "should we own or rent". We welcome conversations long before requirements are settled.