Site Reliability Engineer Hanoi, Vietnam

Reliability,
engineered for
the real world.

I’m Chu Tuan Linh. I design, automate, and operate infrastructure that stays calm under pressure—from global DNS platforms to multi-cloud systems.

Observe Automate Improve
Chu Tuan Linh at the ICANN DNS Symposium in Da Nang
Based in Hanoi UTC +7
Systems mindset See clearly. Act early.
13+ years in infrastructure
~5% of global DNS traffic supported
89% reduction in system alarms
≤30m zero-downtime deployments
Linux / Unix REST APIs Prompt Engineering Local LLMs Terraform Ansible BGP Anycast Prometheus Grafana IBM Cloud AWS Docker

01 About

I turn operational complexity into systems teams can trust.

For more than a decade, I’ve worked where infrastructure, software, and people meet. My job is to make that intersection less fragile.

That means seeing the whole system, automating the repeatable parts, and helping teams make clear decisions when the unexpected happens. I’ve led incident response across departments, supported follow-the-sun operations, and documented complex architectures so knowledge moves with the team.

01

Make it visibleBuild observability before guessing.

02

Make it repeatableAutomate toil and encode good decisions.

03

Make it sharedTurn incidents into durable team knowledge.

02 Expertise

The layers behind dependable systems.

Deep infrastructure experience, paired with the tools and operating habits that keep platforms healthy.

01

Platform reliability

Operating high-availability Linux infrastructure across bare metal, virtual machines, and serverless workloads.

  • Linux / Unix
  • Bare metal
  • DNS
  • BGP Anycast
  • Networking
02

Cloud & automation

Building repeatable infrastructure and safer deployment paths across public, private, and hybrid clouds.

  • Terraform
  • Ansible
  • AWX
  • Argo CD
  • Jenkins
  • IBM Cloud
  • AWS
03

Observability & response

Designing signal-rich monitoring and coordinating calm, cross-functional responses to production problems.

  • Prometheus
  • Grafana
  • ELK
  • Datadog
  • Splunk
  • PagerDuty
04

Systems & data

Connecting applications, data stores, queues, and scripts into maintainable operational systems.

  • Docker
  • PostgreSQL
  • MongoDB
  • Redis
  • RabbitMQ
  • Python
  • Bash
05

New capability

AI systems & integration

Deploying local language models and connecting them to practical workflows through APIs, thoughtful prompt design, and infrastructure automation.

  • REST APIs
  • Prompt engineering
  • Local LLM deployment
  • Python
  • Automation

03 Selected impact

Better systems show up in the numbers.

Technical choices matter most when they create measurable breathing room for users and teams.

89%

Fewer system alarms

Reduced noise so operators could focus on meaningful signals and act faster.

≤30min

Cloud Mesh deployments

Cut deployment time from hours to under 30 minutes, with zero downtime.

99.99%

Infrastructure uptime

Maintained dependable on-premises operations early in my systems career.

04 Experience

Built in production.
Proven under pressure.

2019 — Now

Hanoi · Global

NS1 an IBM company

Site Reliability Engineer

Operating global DNS and multi-cloud infrastructure, improving visibility, automating change, and helping teams resolve complex problems before they become customer incidents.

  • Help manage thousands of bare-metal, VPS, and serverless workloads serving approximately 5% of global DNS queries.
  • Strengthened real-time observability across a wide monitoring stack and coordinated cross-department incident response as Problem Commander.
  • Automated infrastructure and continuous delivery with Terraform, Ansible, AWX, and Argo CD.
Global DNSHybrid cloudSREInfrastructure as Code

2013 — 2019

Hanoi · Global support

Tek Experts

Team Leader, L3 · HPE Customer Support

Led technical support delivery across regions, connecting customers, frontline engineers, and R&D while keeping team knowledge and resources aligned to the SLA.

  • Supported a 24/7 follow-the-sun operation across EMEA, the Americas, and APAC.
  • Reached 4.7/5 customer satisfaction and balanced 90% of team workload.
  • Also supported and customized Microsoft Dynamics 365 and ERP solutions as a senior consultant.
Technical leadershipHP-UX / AIXDynamics 365Customer reliability

2012 — 2013

Hanoi

3S Intersoft JSC

System Administrator

Managed IT assets and on-premises infrastructure, and built the networking and development-tool foundation that kept engineering work moving.

  • Maintained 99.99% infrastructure uptime.
  • Configured switching, firewalls, load balancing, Git/SVN, and Redmine.
  • Recognized as Best Employee of 2012.
Systems administrationNetworkingDeveloper tooling

05 Recognition

Trusted by teams.
Recognized for impact.

The best outcomes are shared. These moments reflect both individual contribution and the teams I helped move forward.

2024

NS1 · IBM

Outstanding Performance

2021

NS1 · IBM

Amazing Contributor

2018

Tek Experts

Best Team Performance

2015

Tek Experts

Outstanding Engineer

06 Foundations

Always learning.
Always sharing.

Away from dashboards, I explore AI, LLMs, blockchain, read science, and stay curious about how the world works.

Education

Bachelor of Computer Science

Hanoi University

Systems & networking studies

Bachkhoa Information Technology Academy

Training & credentials

CCNP Routing

Cisco networking

Linux Site Trainer

Tek Experts

HP BSM Operations Manager

Unix operations

07 Contact

Let’s make systems
calmer.

Have an infrastructure problem, a reliability challenge, or simply a good systems story? I’d be glad to hear it.