09/11/2026

AVP SRE and Cloud Solutions

Job Description

Why GM Financial Technology?

Innovation isn’t just a talking point at GM Financial, it’s how we operate. From generative AI and cloud-native technologies to peer-led learning and hackathons, our tech teams are building real solutions that make a difference. We’re committed to AI-powered transformation, using advanced machine learning and automation to help us reimagine customer interactions and modernize operations, positioning GM Financial as a leader in digital innovation within a dynamic industry. 

Join us and discover a workplace where your ideas matter, your development is prioritized, and you can truly make a global impact.

 

This position will be posted until filled. 

About The Role: 

The AVP, Platform Site Reliability Engineering (PSRE), is a strategic technology and people leader responsible for executing the reliability engineering vision, strategy, and operating model for the organization. This leader advances an AI-enabled reliability organization focused on resilient, self-healing systems, operational excellence, technology risk reduction, and software engineering best practices. 

The AVP establishes standards, governance, and agentic AI capabilities that automate reliability, operational readiness, vulnerability and technology debt remediation, cloud cost optimization, and non-functional requirement validation throughout the software delivery lifecycle. The role develops high-performing teams of Site Reliability Engineers (SREs) and Site Reliability Analysts (SRAs) while fostering innovation, continuous learning, accountability, and toil elimination. 

  • Define and execute the PSRE strategy, roadmap, standards, and operating model aligned with business and technology objectives. 
  • Lead the adoption and responsible governance of agentic AI across reliability engineering, incident management, monitoring, observability, technology risk, vulnerability and technology debt remediation, cloud cost optimization, and non-functional requirements. 
  • Establish automated controls and intelligent agents that validate reliability, security, operational readiness, and non-functional requirements before software progresses through the delivery lifecycle, including before pull requests where practical. 
  • Drive resilient, self-healing systems that detect, diagnose, remediate, and recover from known failure conditions with appropriate safeguards, human oversight, and auditability. 
  • Champion engineering practices that improve availability, scalability, recoverability, performance, maintainability, and operational sustainability while reducing manual intervention and toil. 
  • Establish observability strategies across metrics, logs, traces, application performance, service health, and user-experience signals using agent-driven correlation, anomaly detection, diagnostics, and proactive remediation. 
  • Drive measurable improvements in incident trends, recurring-issue elimination, MTTR, service availability, alert quality, customer impact, and operational efficiency. 
  • Ensure recurring incidents are systematically addressed through root cause remediation, automation, self-healing capabilities, and durable engineering improvements. 
  • Establish operational governance for incident and escalation management, operational readiness, service ownership, documentation, runbooks, supportability, and service health. 
  • Partner with Cybersecurity, Architecture, Product, and Engineering teams to reduce technology risk and improve vulnerability remediation, technology debt, resiliency, and engineering standards. 
  • Establish SLOs, SLIs, error budgets, reliability KPIs, operational metrics, and executive reporting that connect technology performance to customer and business outcomes. 
  • Build and lead high-performing teams by developing engineers, analysts, and technical leaders through coaching, mentorship, structured career development, succession planning, and continuous technical upskilling. 
  • Establish a culture of knowledge sharing, innovation, engineering excellence, accountability, and continuous improvement that strengthens organizational capability and grows future technical and people leaders. 
  • Provide executive leadership during major incidents and influence technology strategy, investment decisions, and engineering priorities across the organization. 

Apply Now