DS DevShelfHub Projects · AI tools
DE
Full time 🤖 AI / ML

DevOps Engineer II (Applied AI Site Reliability Engineer) at Deloitte

Listed by DevShelfHub

India (USI)
Hybrid
Mid-level · 4–8 yrs

About the Role

Deloitte is hiring an Applied AI Site Reliability Engineer II to drive reliability, performance, and operational integrity of high-visibility products and platforms. This role is part of US Deloitte Technology Product Engineering, delivering innovative digital solutions to businesses and internal operations.

This role requires 4+ years of software engineering and SRE experience operating large-scale, distributed, cloud-native systems in production.

What you'll do

  • You will keep production safe, performant, and cost-effective while supporting admission of systems into production.
  • Embrace accountability for reliability, performance, and cost outcomes measured in SLOs and error budgets
  • Serve as technical advocate for production reliability, operability, and graceful degradation
  • Maintain operational integrity of production and pre-production environments
  • Build and operate codified, version-controlled observability dashboards and SLO-driven alerting
  • Run performance, ambient-noise, and chaos testing to verify readiness
  • Guard environments against drift and enforce segregation-of-duties controls
  • Create technical specifications, runbooks, and shared playbooks
  • Lead blameless postmortems that turn incidents into systemic fixes
  • Develop lean operational solutions through rapid experimentation
  • Co-define service-level objectives and verify operational readiness with product teams
  • Operate AI/ML and agentic workloads reliably including drift, train/serve skew, and cost anomalies

What we're looking for

  • Bachelor's degree in computer science, software engineering, data science, machine learning, or related discipline
  • 4+ years software engineering and SRE experience operating large-scale, distributed, cloud-native systems in production
  • Experience with Python, Go, Bash, Java, C#/.NET, SQL/NoSQL, Kubernetes, Terraform, ArgoCD, CI/CD, and observability stacks
  • 2+ years SRE or production engineering: SLIs, SLOs, SLAs, error budgets, incident command, on-call

Skills & Technologies

Required

Bachelor's degree in computer science, software engineering, data science, machine learning, or related discipline 4+ years software engineering and SRE experience operating large-scale, distributed, cloud-native systems in production Experience with Python, Go, Bash, Java, C#/.NET, SQL/NoSQL, Kubernetes, Terraform, ArgoCD, CI/CD, and observability stacks 2+ years SRE or production engineering: SLIs, SLOs, SLAs, error budgets, incident command, on-call Production observability: OpenTelemetry, Prometheus, Grafana, Datadog, CloudWatch, Azure Monitor, Splunk 2+ years cloud-native engineering on Azure, AWS, or GCP including AI/ML services Operating AI/ML and agentic workloads in production with MLOps/LLMOps experience Load and performance testing (k6, JMeter), chaos engineering, capacity planning, FinOps

Benefits & Perks

  • Health Insurance
  • Flexible Leave Policy
  • Learning Budget
  • EPF / NPS

Why This Role is Good for Experienced Professionals

  • This role offers SRE leadership for AI and agentic workloads alongside cloud-native platform engineering at global scale.
  • Drive culture of accountability for reliability, performance, and cost through SLOs and error budgets
  • Serve as technical advocate for production reliability and operability
  • Build and operate production observability with SLO-driven alerting
  • Run performance, ambient-noise, and chaos testing to verify readiness
  • Operate AI/ML and agentic workloads including MLOps/LLMOps failure modes
  • Apply cloud platform ownership on Azure, AWS, or GCP including AI/ML services
  • Lead blameless postmortems and write high-quality reliability automation
  • Co-define service-level objectives with engineering teams
  • Hands-on infrastructure-focused engineering with continuous learning
  • Collaborate with platform engineering, security, and product teams

About Deloitte

US Deloitte Technology Product Engineering has modernized software and product delivery, creating a scalable, cost-effective model focused on value and outcomes. Product Engineering delivers innovative digital solutions and serves as the engine that drives Deloitte's success.
Industry
Technology
Company Size
500+
Website
deloitte.com
🚀 Boost your chances

Skip the queue

Contact the recruiter directly at Deloitte and follow up personally — most applicants never do.

Browse Recruiter Database

Free trial · 40 contacts · ₹0

Job Details

Type
Full time
Level
Mid-level
Experience
4–8 years
Location
India (USI)
Work Mode
Hybrid
Category
AI / ML

Share this Job

Recruiter Database

Want to stand out? Talk to the recruiter directly.

Most applicants never hear back. Our Recruiter Database gives you direct email access to the people making hiring decisions — so you can follow up personally and actually get noticed.

  • Recruiter name, company & direct email address
  • Sourced from active job postings across top companies
  • Free trial — 40 contacts, no credit card needed

₹0

to get started

Get Free Trial