Site Reliability Engineer Hanoi, Vietnam
Reliability,
engineered for
the real world.
I’m Chu Tuan Linh. I design, automate, and operate infrastructure that stays calm under pressure—from global DNS platforms to multi-cloud systems.
01 About
I turn operational complexity into systems teams can trust.
For more than a decade, I’ve worked where infrastructure, software, and people meet. My job is to make that intersection less fragile.
That means seeing the whole system, automating the repeatable parts, and helping teams make clear decisions when the unexpected happens. I’ve led incident response across departments, supported follow-the-sun operations, and documented complex architectures so knowledge moves with the team.
Make it visibleBuild observability before guessing.
Make it repeatableAutomate toil and encode good decisions.
Make it sharedTurn incidents into durable team knowledge.
02 Expertise
The layers behind dependable systems.
Deep infrastructure experience, paired with the tools and operating habits that keep platforms healthy.
Platform reliability
Operating high-availability Linux infrastructure across bare metal, virtual machines, and serverless workloads.
Cloud & automation
Building repeatable infrastructure and safer deployment paths across public, private, and hybrid clouds.
Observability & response
Designing signal-rich monitoring and coordinating calm, cross-functional responses to production problems.
Systems & data
Connecting applications, data stores, queues, and scripts into maintainable operational systems.
New capability
AI systems & integration
Deploying local language models and connecting them to practical workflows through APIs, thoughtful prompt design, and infrastructure automation.
03 Selected impact
Better systems show up in the numbers.
Technical choices matter most when they create measurable breathing room for users and teams.
Fewer system alarms
Reduced noise so operators could focus on meaningful signals and act faster.
Cloud Mesh deployments
Cut deployment time from hours to under 30 minutes, with zero downtime.
Infrastructure uptime
Maintained dependable on-premises operations early in my systems career.
04 Experience
Built in production.
Proven under pressure.
NS1 an IBM company
Site Reliability Engineer
Operating global DNS and multi-cloud infrastructure, improving visibility, automating change, and helping teams resolve complex problems before they become customer incidents.
- Help manage thousands of bare-metal, VPS, and serverless workloads serving approximately 5% of global DNS queries.
- Strengthened real-time observability across a wide monitoring stack and coordinated cross-department incident response as Problem Commander.
- Automated infrastructure and continuous delivery with Terraform, Ansible, AWX, and Argo CD.
Tek Experts
Team Leader, L3 · HPE Customer Support
Led technical support delivery across regions, connecting customers, frontline engineers, and R&D while keeping team knowledge and resources aligned to the SLA.
- Supported a 24/7 follow-the-sun operation across EMEA, the Americas, and APAC.
- Reached 4.7/5 customer satisfaction and balanced 90% of team workload.
- Also supported and customized Microsoft Dynamics 365 and ERP solutions as a senior consultant.
3S Intersoft JSC
System Administrator
Managed IT assets and on-premises infrastructure, and built the networking and development-tool foundation that kept engineering work moving.
- Maintained 99.99% infrastructure uptime.
- Configured switching, firewalls, load balancing, Git/SVN, and Redmine.
- Recognized as Best Employee of 2012.
05 Recognition
Trusted by teams.
Recognized for impact.
The best outcomes are shared. These moments reflect both individual contribution and the teams I helped move forward.
NS1 · IBM
Outstanding Performance
NS1 · IBM
Amazing Contributor
Tek Experts
Best Team Performance
Tek Experts
Outstanding Engineer
06 Foundations
Always learning.
Always sharing.
Away from dashboards, I explore AI, LLMs, blockchain, read science, and stay curious about how the world works.
Education
Bachelor of Computer Science
Hanoi University
Systems & networking studies
Bachkhoa Information Technology Academy
Training & credentials
CCNP Routing
Cisco networking
Linux Site Trainer
Tek Experts
HP BSM Operations Manager
Unix operations
07 Contact
Let’s make systems
calmer.
Have an infrastructure problem, a reliability challenge, or simply a good systems story? I’d be glad to hear it.