Site reliability · AI/ML infrastructure
Reliability, at depth.
Site reliability engineer for AI/ML platforms — building and operating GPU fleets, distributed training pipelines, and large-scale data infrastructure for organizations from research institutes to global trading platforms.
live — the pod behind this page is a Rust particle simulation compiled to wasm32, ~1,900 dots ticking in linear memory.
§ Sonar
Query this CV
This page's data lives in a DuckDB database running in your browser via
WebAssembly. Run SQL against it — an UPDATE re-renders the
dive record below.
§ About
SRE practice, ML background
I work where production infrastructure meets machine learning: Kubernetes, Terraform, observability, and networking on one side; a graduate research background and hands-on MLOps on the other. That combination is what it takes to run GPU platforms and training pipelines that researchers can rely on — and to retire the toil that keeps them from scaling.
- CKA certified
- Kaggle Competitions Expert
- OSS: Ray · DeepEval · Loki Helm chart
- K8s / Terraform / GPU infra
§ Dive record
Experience, by descent
Most recent at the surface — each stop one dive deeper into the stack.
-
▾ 0 m — present
Site Reliability Engineer III · F5
- Shipped an AppStack feature that installs NVIDIA drivers and CUDA for modern GPUs across managed Kubernetes clusters.
- Enabled deployment on KVM-class environments with L4 GPU passthrough via AppStack VMs.
- Took over a stalled Loki logging-platform migration and led it to completion — on-premise to EKS with an S3 backend, eliminating hardware-capacity toil and enabling on-demand scale-out.
- Retired the in-house templating system behind the Loki deployment in favor of the official Helm chart, aligning operations with community standards.
-
▾ 150 m
MLOps Engineer · Cluster (freelance)
- Tuned training-data load performance on an HPC system for ML using striping — accelerating training by ~30%.
- Built a distributed training pipeline with Ray and Weights & Biases on AWS.
- Drove an LLM chatbot project — evaluating models, methods, and vector DBs — then handed the system to SWEs.
-
▾ 300 m —
DevOps Engineer · Binance
- Led development of a GitOps CI/CD pipeline for Terraform from scratch, integrating Atlantis and Policy-as-Code with OPA.
- Directed a VPN migration across multiple AWS accounts using Transit Gateway, VPC endpoints, and Client VPN.
- Revived legacy petabyte-scale Hive data pipelines — automated, visualized, and documented them with backfill and retry in a semi-idempotent design.
-
▾ 450 m —
Co-Founder / Tech Lead · MaskTap
demo video ↗- Designed and built an auto-anonymization product from scratch around an extensible API.
- Developed a framework that automatically masks PII in compliance with GDPR and APPI — original enough to be considered for a patent application.
-
▾ 600 m —
Machine Learning Engineer · Research Center for Advanced Science and Technology, The University of Tokyo (freelance)
arXiv paper ↗- Created a computer-vision segmentation model end to end, starting from planning what data to collect and how much.
- Diagnosed an unbalanced, undersized dataset in the first PoC and delivered a model that generalizes — the project resulted in an arXiv publication.
-
▾ 750 m —
Site Reliability Engineer · SoftBank
- Built and operated a production Kubernetes ML-inference platform, with scheduled autoscaling that absorbed traffic surges during large events.
- Designed end-user APIs on the AI platform using Kubernetes CRDs while preserving compatibility for existing users.
- Shipped a serverless Vue web app designed for easy maintenance and scale.
-
▾ 900 m —
ML Engineer / Tech Lead · BEENOS (part-time → freelance)
- Led a team of five as tech lead: Python, Node, AWS, Terraform, GitHub Actions, Kubernetes.
- Built an ML system for importance discrimination end to end — data collection through deployment.
- Replaced a legacy monolith crawler with a Pub/Sub architecture for reliability.
§ Instruments
Working range
SRE
- Kubernetes
- CI/CD
- AWS
- Google Cloud
- Terraform
- Linux
- Networking
Backend
- Python
- Go
- Rust
- Java
- FastAPI
- Django
- Laravel
Machine learning
- PyTorch
- Deep learning
- NLP
- MLOps
- Prompt engineering
Observability
- Loki
- Prometheus
- Grafana
- Vector
- Fluent Bit
- Elasticsearch
Data
- Hive
- Spark
- Kafka
- ClickHouse
- BigQuery
- Dagster
§ Signals
Beyond the day job
Open source
- Contributor to Ray, DeepEval, the Loki Helm chart (Grafana community), and official AWS repositories upstream fixes and features grown out of production SRE and MLOps work
Awards
- Top 3 — PWS Cup 2021, privacy-protection technology competition Computer Security Symposium · 2021
Certifications
- CKA — Certified Kubernetes Administrator The Linux Foundation · 2024
- Kaggle Competitions Expert Kaggle · 2024
- AWS Certified SysOps Administrator — Associate AWS · 2021
- On-The-Ground Special Radio Operator The Government of Japan · 2018
- JSBA Badge Test Level 1 Japan Snowboarding Association · 2026