[Remote] Sr Machine Learning Engineering Manager - AI Quality and Governance
Auto ImportNote: The job is a remote job and is open to candidates in USA. Workiva is a platform designed to bring confidence, control, and a competitive edge to complex organizations. The Sr Machine Learning Engineering Manager - AI Quality and Governance will lead a multidisciplinary team to establish quality standards and governance for AI products, ensuring they are trustworthy and ready for enterprise use.
Responsibilities
- Lead, mentor, and develop a multidisciplinary team of software, ML, and quality engineers
- Build a culture of technical excellence, quality ownership, experimentation, and continuous improvement
- Establish clear team priorities while balancing platform investments, product needs, and enterprise risk
- Recruit engineers with complementary expertise across software quality, ML evaluation, platform engineering, and governance automation
- Define and drive a comprehensive quality strategy for Workiva's AI platform and products, spanning unit, integration, end-to-end, performance, resilience, security, and production testing
- Establish measurable quality bars, release-readiness criteria, and automated quality gates for AI and agentic capabilities
- Advance testing approaches for nondeterministic systems, including RAG pipelines, agents, prompts, models, tools, and multi-step workflows
- Detect regressions, model or data drift, unsafe behavior, and degraded customer experiences before and after release
- Lead architecture and delivery of a scalable, self-service evaluation platform for generative AI, RAG, and agentic systems
- Enable teams to create, manage, version, and reuse evaluation datasets, golden test sets, task-specific metrics, graders, and benchmarks
- Support deterministic checks, statistical metrics, model-based graders, human evaluation, adversarial testing, and domain-expert review
- Build capabilities for offline evaluation, pre-release regression testing, online experimentation, production sampling, and continuous evaluation
- Ensure evaluation results are reproducible, explainable, actionable, and integrated into developer workflows, CI/CD pipelines, and operational dashboards
- Translate Workiva's Responsible AI principles into practical engineering controls and platform capabilities
- Build governance into the AI lifecycle through traceability, lineage, versioning, documentation, risk classification, approval workflows, and auditable evidence
- Partner with Security, Legal, Privacy, Compliance, and Risk teams to define controls that support enterprise and regulated use cases
- Enable inventories and traceability across models, prompts, datasets, evaluations, tools, knowledge sources, and deployed AI features
- Collaborate with Product, Program Management, UX, UXR, Data Science, Security, Legal, Risk, and engineering leaders to define quality expectations and roadmaps
- Influence engineering teams across Workiva to adopt shared evaluation standards, testing practices, observability, and release controls
- Communicate complex technical tradeoffs, quality signals, and risk findings clearly to technical and non-technical audiences
- Ensure the evaluation and governance platform is secure, scalable, reliable, observable, and cost-effective
- Define service-level objectives and meaningful operational and quality metrics
- Champion production readiness, incident response, root-cause analysis, and continuous operational improvement
Skills
- Bachelor's degree in Computer Science, Engineering, Data Science, or related field (or equivalent experience)
- 10+ years in software engineering, ML engineering, quality engineering, or related roles, including 4+ years leading an engineering team
- Strong software engineering and systems-design fundamentals, with experience delivering and operating production SaaS or platform capabilities
- Demonstrated experience establishing automated quality practices for distributed, cloud-based products
- Practical understanding of the generative AI development lifecycle and challenges of evaluating nondeterministic systems
- Experience with generative AI concepts: LLMs, RAG, embeddings, vector/hybrid search, agents, tool use, and prompt orchestration
- Experience defining measurable quality criteria using data, experimentation, telemetry, and production signals
- Experience with cloud-native architectures on AWS, Azure, or GCP
- Proven ability to lead senior individual contributors, navigate technical disagreements, and build high-performance cultures
- Strong communication and cross-functional leadership skills
- Master's degree in Computer Science, Engineering, ML, Data Science, or related field
- Experience building or operating AI/ML evaluation, experimentation, observability, model-governance, or ML platform capabilities
- Experience evaluating RAG and agentic systems, including retrieval quality, groundedness, task completion, tool use, and safety
- Familiarity with evaluation techniques: golden datasets, statistical metrics, model-based graders, human evaluation, red teaming, A/B testing, and drift/regression detection
- Working knowledge of ML/AI lifecycle practices: dataset management, model/prompt versioning, experiment tracking, deployment, monitoring, and feedback loops
- Experience translating Responsible AI, model-risk, privacy, security, or regulatory requirements into scalable engineering controls
- Familiarity with AI risk/governance frameworks (NIST AI RMF, ISO/IEC 42001, or comparable)
- Experience with Kubernetes, microservices, CI/CD, infrastructure as code, and modern DevOps/MLOps practices
- Experience supporting enterprise software in regulated or high-assurance environments
Benefits
- A discretionary bonus typically paid annually
- Restricted Stock Units granted at time of hire
- 401(k) match
- Comprehensive employee benefits package
- Workiva supports employees in working where they work best - either from an office or remotely from any location within their country of employment.
Company Overview
Company H1B Sponsorship