depth.log

Site reliability · AI/ML infrastructure

Reliability, at depth.

Site reliability engineer for AI/ML platforms — building and operating GPU fleets, distributed training pipelines, and large-scale data infrastructure for organizations from research institutes to global trading platforms.

live — the pod behind this page is a Rust particle simulation compiled to wasm32, ~1,900 dots ticking in linear memory.

§ About

SRE practice, ML background

I work where production infrastructure meets machine learning: Kubernetes, Terraform, observability, and networking on one side; a graduate research background and hands-on MLOps on the other. That combination is what it takes to run GPU platforms and training pipelines that researchers can rely on — and to retire the toil that keeps them from scaling.

§ Dive record

Experience, by descent

Most recent at the surface — each stop one dive deeper into the stack.

  1. ▾ 0 m — present

    Site Reliability Engineer III · F5

    • Shipped an AppStack feature that installs NVIDIA drivers and CUDA for modern GPUs across managed Kubernetes clusters.
    • Enabled deployment on KVM-class environments with L4 GPU passthrough via AppStack VMs.
    • Took over a stalled Loki logging-platform migration and led it to completion — on-premise to EKS with an S3 backend, eliminating hardware-capacity toil and enabling on-demand scale-out.
    • Retired the in-house templating system behind the Loki deployment in favor of the official Helm chart, aligning operations with community standards.
  2. ▾ 150 m

    MLOps Engineer · Cluster (freelance)

    • Tuned training-data load performance on an HPC system for ML using striping — accelerating training by ~30%.
    • Built a distributed training pipeline with Ray and Weights & Biases on AWS.
    • Drove an LLM chatbot project — evaluating models, methods, and vector DBs — then handed the system to SWEs.
  3. ▾ 300 m

    DevOps Engineer · Binance

    • Led development of a GitOps CI/CD pipeline for Terraform from scratch, integrating Atlantis and Policy-as-Code with OPA.
    • Directed a VPN migration across multiple AWS accounts using Transit Gateway, VPC endpoints, and Client VPN.
    • Revived legacy petabyte-scale Hive data pipelines — automated, visualized, and documented them with backfill and retry in a semi-idempotent design.
  4. ▾ 450 m

    Co-Founder / Tech Lead · MaskTap

    demo video ↗
    • Designed and built an auto-anonymization product from scratch around an extensible API.
    • Developed a framework that automatically masks PII in compliance with GDPR and APPI — original enough to be considered for a patent application.
  5. ▾ 600 m

    Machine Learning Engineer · Research Center for Advanced Science and Technology, The University of Tokyo (freelance)

    arXiv paper ↗
    • Created a computer-vision segmentation model end to end, starting from planning what data to collect and how much.
    • Diagnosed an unbalanced, undersized dataset in the first PoC and delivered a model that generalizes — the project resulted in an arXiv publication.
  6. ▾ 750 m

    Site Reliability Engineer · SoftBank

    • Built and operated a production Kubernetes ML-inference platform, with scheduled autoscaling that absorbed traffic surges during large events.
    • Designed end-user APIs on the AI platform using Kubernetes CRDs while preserving compatibility for existing users.
    • Shipped a serverless Vue web app designed for easy maintenance and scale.
  7. ▾ 900 m

    ML Engineer / Tech Lead · BEENOS (part-time → freelance)

    • Led a team of five as tech lead: Python, Node, AWS, Terraform, GitHub Actions, Kubernetes.
    • Built an ML system for importance discrimination end to end — data collection through deployment.
    • Replaced a legacy monolith crawler with a Pub/Sub architecture for reliability.

§ Instruments

Working range

SRE

  • Kubernetes
  • CI/CD
  • AWS
  • Google Cloud
  • Terraform
  • Linux
  • Networking

Backend

  • Python
  • Go
  • Rust
  • Java
  • FastAPI
  • Django
  • Laravel

Machine learning

  • PyTorch
  • Deep learning
  • NLP
  • MLOps
  • Prompt engineering

Observability

  • Loki
  • Prometheus
  • Grafana
  • Vector
  • Fluent Bit
  • Elasticsearch

Data

  • Hive
  • Spark
  • Kafka
  • ClickHouse
  • BigQuery
  • Dagster

§ Signals

Beyond the day job

Open source

  • Contributor to Ray, DeepEval, the Loki Helm chart (Grafana community), and official AWS repositories upstream fixes and features grown out of production SRE and MLOps work

Awards

  • Top 3 — PWS Cup 2021, privacy-protection technology competition Computer Security Symposium · 2021

Certifications

  • CKA — Certified Kubernetes Administrator The Linux Foundation · 2024
  • Kaggle Competitions Expert Kaggle · 2024
  • AWS Certified SysOps Administrator — Associate AWS · 2021
  • On-The-Ground Special Radio Operator The Government of Japan · 2018
  • JSBA Badge Test Level 1 Japan Snowboarding Association · 2026