Staff Site Reliability Engineer
Auto Import<p style="min-height:1.5em"><strong>Build the Future Workforce</strong></p><p style="min-height:1.5em">Wand turns AI into labor. It enables humans and AI agents to operate together as a unified, hybrid workforce, with comprehensive management and oversight. And it’s already operating at scale inside some of the world’s largest organizations.</p><p style="min-height:1.5em">Wand built the world’s first Agentic Labor Infrastructure enabling governments and global enterprises to create, manage, and scale digital workforces.</p><p style="min-height:1.5em">Our mission is to integrate agent ecosystems into the core of work and business, unlocking a generational leap in the global economy. We’re building the infrastructure that lets humans and AI agents operate together safely, transparently, and at scale.</p><p style="min-height:1.5em"><strong>Join Wand in leading the Agentic Shift</strong></p><p style="min-height:1.5em">Wand is building a high-performing global team who take full ownership of what they build. We lead by example, move fast, make data-aware decisions, and continuously push for more- always with a focus on delivering real value to customers.</p><p style="min-height:1.5em">You would be joining a world-class team that combines deep research expertise and real-world product execution, with experience spanning Deepmind, Google, Amazon, Miro, Elise AI, IBM and Accern.</p><p style="min-height:1.5em"></p><p style="min-height:1.5em"><strong>Position Summary</strong></p><p style="min-height:1.5em">We are hiring for a highly experienced Senior Staff SRE Engineer to act as a senior technical authority within our reliability function.</p><p style="min-height:1.5em">This is a deeply hands-on individual contributor role, to build and operate SRE practices at scale. You will design and evolve resilient infrastructure, drive reliability across multiple engineering streams, and ensure our AI-driven products operate with high availability, performance, and security.</p><p style="min-height:1.5em">You will work across platform, product, data, and ML teams, helping us productionise models, absorb and standardise customer environments, strengthen Kubernetes-based architecture, and mature our CI/CD pipelines end-to-end.</p><p style="min-height:1.5em">You will also collaborate with other Staff engineers and Architects to shape the global product architect and technology vision.</p><p style="min-height:1.5em"></p><p style="min-height:1.5em"><strong>Responsibilities</strong></p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">Architect, deploy, and operate scalable, secure production environments (AWS preferred).</p></li><li><p style="min-height:1.5em">Lead reliability improvements across multiple engineering streams.</p></li><li><p style="min-height:1.5em">Design and evolve Kubernetes-based infrastructure, including migration and optimisation initiatives.</p></li><li><p style="min-height:1.5em">Build and enforce strong Infrastructure-as-Code standards.</p></li><li><p style="min-height:1.5em">Define and operationalise SLIs, SLOs, and error budgets.</p></li><li><p style="min-height:1.5em">Strengthen observability across applications, infrastructure, data pipelines, and ML systems.</p></li><li><p style="min-height:1.5em">Work closely with product and data teams to integrate model analytics and product telemetry into reliability insights.</p></li><li><p style="min-height:1.5em">Work across and optimise the entire CI/CD pipeline, from build to deploy to rollback.</p></li><li><p style="min-height:1.5em">Improve release safety, deployment frequency, and predictability of SLAs.</p></li><li><p style="min-height:1.5em">Lead incident response for complex cross-system failures and drive postmortems.</p></li><li><p style="min-height:1.5em">Reduce operational toil through automation and platform engineering improvements.</p></li><li><p style="min-height:1.5em">Design processes and tooling to absorb, standardise, and troubleshoot customer environments.</p></li><li><p style="min-height:1.5em">Support and productionise ML workloads (MLOps practices including model deployment, monitoring, retraining workflows).</p></li><li><p style="min-height:1.5em">Ensure infrastructure aligns with enterprise-grade security and regulatory requirements.</p></li><li><p style="min-height:1.5em">Mentor engineers and raise the overall reliability bar across teams.</p></li></ul><p style="min-height:1.5em"></p><p style="min-height:1.5em"><strong>Key Requirements</strong></p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">Extensive hands-on experience in SRE or Production Engineering roles.</p></li><li><p style="min-height:1.5em">Demonstrated experience building or scaling SRE practices in high-growth or complex environments.</p></li><li><p style="min-height:1.5em">Deep expertise in AWS or Azure-based cloud infrastructure.</p></li><li><p style="min-height:1.5em">Strong experience with Kubernetes (including migration, scaling, and production hardening).</p></li><li><p style="min-height:1.5em">Advanced Infrastructure-as-Code experience (Terraform or equivalent).</p></li><li><p style="min-height:1.5em">End-to-end CI/CD pipeline design and optimisation experience.</p></li><li><p style="min-height:1.5em">Strong experience with observability tooling across distributed systems.</p></li><li><p style="min-height:1.5em">Experience troubleshooting complex multi-tenant or customer-hosted environments.</p></li><li><p style="min-height:1.5em">Experience supporting production data platforms and ML systems.</p></li><li><p style="min-height:1.5em">MLOps experience, including model deployment and monitoring.</p></li><li><p style="min-height:1.5em">Strong understanding of distributed systems, scalability, and fault tolerance.</p></li><li><p style="min-height:1.5em">Systems thinker who understands interactions across infrastructure, product, data, and ML.</p></li><li><p style="min-height:1.5em">Excellent communication skills and ability to work cross-functionally.</p></li></ul><p style="min-height:1.5em"></p><p style="min-height:1.5em"><strong>Preferred Experience</strong></p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">Experience in large-scale global B2B/B2C products.</p></li><li><p style="min-height:1.5em">Experience working with AI/ML systems, NLP, or LLM-based products.</p></li><li><p style="min-height:1.5em">Experience integrating product analytics and model performance metrics into operational monitoring.</p></li><li><p style="min-height:1.5em">Background in enterprise environments with strong security and compliance requirements.</p></li><li><p style="min-height:1.5em">Experience implementing regulatory controls within cloud infrastructure.</p></li><li><p style="min-height:1.5em">Experience scaling infrastructure during rapid growth phases.</p></li><li><p style="min-height:1.5em">Experience evaluating infrastructure tooling and vendors.</p></li><li><p style="min-height:1.5em">Experience in collaborating with large scale enterprise customers to deploy and operate environments within their accounts and VPCs.</p></li></ul><p style="min-height:1.5em"></p><p style="min-height:1.5em"><strong>Personal Characteristics</strong></p><ul style="min-height:1.5em"><li><p style="min-height:1.5em">Strong problem solver who anticipates failure modes.</p></li><li><p style="min-height:1.5em">High ownership mentality and accountability.</p></li><li><p style="min-height:1.5em">Comfortable working across streams and influencing without formal authority.</p></li><li><p style="min-height:1.5em">Learning-oriented with a drive for continuous improvement.</p></li></ul>