Anonymized case · business-critical workflow
Available for selected consulting, contract & fractional work
Tai VONG
Backend & Platform Engineer · Technical Lead
I help product teams diagnose difficult backend failures, stabilize Go services, and turn fragile systems into observable, production-ready platforms.
- Go & gRPC
- Distributed systems
- Production reliability
- Technical leadership
SYSTEM ONLINE 10.7769°N 106.7009°E Ho Chi Minh City
- 01Signal
- 02Isolate
- 03Fix
- 04Verify
Selected work
Open source, writing, and production case work
Research · computer vision
Maintaining face trajectories through fragmented detections
Case study · gRPC platform tooling
Ad hoc gRPC boilerplate repeated across every new service
Ways to work together
From a five-day diagnostic to ongoing technical direction
Backend Reliability Diagnostic
Focused 5-day investigation
A focused five-day investigation into a backend system that's failing in ways nobody can quite explain. You get a failure map, prioritized findings, a concrete remediation plan, and a handoff your team can execute without me.
- Failure map of the system
- Prioritized findings, ranked by risk and effort
- Concrete remediation plan
- Handoff walkthrough with your engineers
Go & gRPC Production Hardening
2–4 week sprint
Takes a Go or gRPC service from "it mostly works" to production-hardened — correctness, observability, failure handling, and operability.
- Correctness & edge-case coverage
- Observability wired in — metrics, tracing, logs
- Failure handling & graceful degradation
- Operability handoff docs & runbooks
Fractional Technical Lead
Part-time, ongoing
Part-time technical direction for teams that need senior judgment without a full-time hire.
- Architecture & technical decisions
- Delivery systems & engineering process
- Mentoring & team growth
Also available for selected contract engagements, technical advisory, speaking & workshops, and expert calls — reach out with what you're working on.
How I work
The same loop, every time
- 01
Observe
Read the signal before touching code — logs, traces, metrics, and the shape of the failure.
- 02
Reproduce
Recreate the failure on demand. If it can't be reproduced, it can't be trusted as fixed.
- 03
Find the invariant
Identify the assumption the system is silently violating.
- 04
Fix the system
Fix the invariant, not just the symptom, so the failure class doesn't come back.
- 05
Verify in production
Confirm the fix holds under real traffic, not just in a test environment.
Writing & open source
Notes on Go, gRPC, and distributed systems
Medium
Implementing common Go HTTP middleware
Middleware as a wrapper that runs before and after the main handler, keeping business logic free of cross-cutting concerns.
Read the articleMedium
Using buf.build to generate your gRPC codes
A tour through migrating an old protobuf/gRPC codegen pipeline to buf, based on hands-on research and practice.
Read the articleResearch paper
A high-performance multi-face tracking system
Detection-tracking and tracklet association in a semi-online framework designed to run at real-time speed.
Read the paperAbout
Tai VONG
Technical Manager at Gearment, leading engineering across product and platform domains — architecture, delivery, and the teams that ship both.
My background is hands-on: backend and platform engineering, Go and gRPC, distributed systems, and the operational work of keeping production systems observable and reliable.
Earlier research background in AI and computer vision occasionally shows up in how I approach problems, though it isn't the focus of my work today.
Ho Chi Minh City, Vietnam
Have a backend problem worth a second opinion?
Tell me what's failing and what you've already tried. I'll reply with whether I think I can help.