My client is a global top tier proprietary trading firm focused on algorithmic and high-frequency trading across global markets. The culture is built for high-agency technologists—genuinely flat structure, fast decision-making, and a “build to win” mentality. They are as much a tech firm as a financial firm, with a strong focus on solving hard problems with smart people. The tech stack includes C++, Rust, Python, Golang, Linux, Kubernetes, and more. They offer exceptional compensation and benefits.

Role Overview

We are seeking a C++ Site Reliability Engineer to join the APAC Production Infrastructure organization (approximately 35 people) in Singapore. This role serves as a critical bridge between core development teams and the production engineering team. You will be responsible for the reliability and uptime of trading infrastructure, handling incident response, root cause analysis, post-mortems, and driving continuous improvement. This is not a traditional ops role—you will need strong Linux skills and C++ at a developer level, as you will be expected to modify C++ code where needed (all trading applications are in C++ on Linux). Open to candidates from any tech background (not necessarily finance), as the HFT domain can be learned. SRE/DevOps experience is a plus but not required—any backend C++ developer open to support responsibilities would be a strong fit. Strong English communication skills are essential.

Key Responsibilities:

  • Technical Expertise: Develop deep technical expertise in assigned product areas and tech stack.
  • Production Ownership: Own production deployment, configuration, and release processes.
  • Performance & Reliability: Drive performance, reliability, and operability through continuous improvement.
  • Tooling: Build and maintain production tooling for deployment, orchestration, monitoring, and diagnostics.
  • Observability: Define and maintain observability, SLI/SLOs, and performance metrics in partnership with product owners.
  • Capacity Planning: Leverage metrics and capacity planning to ensure scalability and uptime.
  • Incident Response: Lead and coordinate incident response, root cause analysis, and post-mortems.
  • Influence Architecture: Promote best practices and influence architecture by aligning with global SRE teams.
  • Documentation & Mentorship: Document processes and procedures; provide mentorship and cross-training to peers.
  • Operational Risk: Actively manage operational risk for production changes.

Required Skills & Experience:

  • Experience: 2–8 years in a backend C++ development role or IT ops role (DevOps, SRE, Linux Systems Engineering, or Network Engineering).
  • C++: Expert-level proficiency in C++—able to read, understand, and modify C++ code as needed.
  • Linux: Strong understanding of Linux OS, including network/system configuration, kernel internals, scheduling, and performance tuning.
  • Networking: Strong understanding of networking concepts such as routing, multicast, LLDP, VLANs, and Ethernet.
  • Communication: Strong English communication skills—able to articulate technical issues clearly and collaborate across global teams.
  • Ownership: Deep sense of ownership and desire to meet business priorities with urgency.
  • On-Call: Ability to handle shared operational and periodic on-call duties.
  • Education: Degree in Computer Science, a related field, or equivalent professional experience.

Please send your CV to Sarah Fan at sarah.fan@ashford-benjamin.com, or call +852 2315 9512 for a confidential discussion.

To apply for this job email your details to sarah.fan@ashford-benjamin.com