Senior Systems Engineer - Observability & Resilience
About the Team
We're refining our Observability and Resilience team that architects and scales intelligent observability patterns across Cox Automotive; enables automated resilence. This isn't traditional monitoring—we create, curate and govern patterns that make it efficient for our engineering organizations to be observable and resilient. To that end we write code that surfaces problems before they become outages, build AI-driven detection and response systems, and enable rapid resolution. This is a ground-floor opportunity to transform the team from implementers of monitoring into technology leaders, true experts in their trade with deep relationship with our engineering organizations.
About the Position
We are seeking highly skilled system engineers for the observability and resilience team. Individuals should be well versed in a variety of technical infrastructure and engineering capabilities and keenly focused on understanding and improving observability and resilience. We are focused on AI, automation and determined to continuously drive efficiencies in our incident response.
This individual should be focused on establishing spec driven standards for operations generally and for observability specifically. You will form strong relationship with our engineering teams and build, curate and champion effective patterns for observability adoption.
We currently work with Cloudwatch, New Relic, Solarwinds, Splunk and rely heavily on Pager Duty and Service Now ITOM for ingestion. We are not seeking users of these tools. We are seeking thought leaders, owners, developers, enablers – emphasis on back end, adoption and user experience.
It is our vision to transform the team from implementers of monitoring to technology leaders; experts in their trade. We will be experts in our tools and drive innovation and adoption.
This role will deliver new capabilities, using AI to develop new tools and new ways to seamlessly adopt the enterprise toolset. You will champion new tools and methods. You will develop and own relationships with engineering teams. You will ensure that our key assets are fully observable and you will help develop automated responses to common issues driving down the time to notice an issue, and crushing the time between recognition and resolution using innovative capabilities.
You will be in rapidly changing, uncertain conditions. AI is changing everything. You must thrive in a challenging and ever evolving environment.
We are in the greatest time ever to be a serious, battle-hardened operator. For the first time, we can specify how the operations harness is to be used while software is being developed. If you are excited about that and seeking a challenge creating, enabling and championing world class observability and response, bring it!!!
We are seeking a skilled and detail-oriented IT Monitoring Engineer to join our team. In this role, you will leverage tools such as ServiceNow ITOM, SolarWinds, and New Relic to ensure seamless monitoring and automation of our IT infrastructure. You will be responsible for administering monitoring tools, writing synthetic testing scripts, automating alert responses, and documenting monitoring solutions. This is a critical role that requires strong technical expertise, communication skills, and the ability to collaborate with cross-functional teams.
What You'll Do:
- Deliver new capabilities, using AI to develop agentic flows and new ways to seamlessly adopt the enterprise toolset
- Champion new tools and methods across the organization
- Develop and own relationships with engineering teams
- Ensure that our key assets are fully observable
- Build automated responses to common issues, driving down the time to notice an issue and crushing the time between recognition and resolution
- Establish spec-driven standards for operations and observability
- Reimagine how we detect, triage, and respond to incidents at scale
- Experiment with new approaches to observability, monitoring, and alerting
- Define what modern observability and response engineering looks like for our organization
Who You Are:
- Bachelor’s degree in a related discipline and 4 years’ experience in a related field. The right candidate could also have a different combination, such as a master’s degree and 2 years’ experience; a Ph.D. and up to 1 year of experience; or 16 years’ experience in a related field
- Focus on Cloud observability and resilience using cloudwatch, New Relic, Service Now ITOM
- Professional experience implementing observability and resilience practices in AWS
- Hands-on with enterprise tooling such as CloudWatch, New Relic, SolarWinds, Splunk, PagerDuty, and ServiceNow ITOM—as a builder and integrator
- Professional experience with writing synthetic tests in Python, Ruby or JavaScript for Playwright Puppeteer or Selenium.
- Distributed systems expertise and understanding of failure modes
- Deep observability experience—instrumentation, metrics, logs, traces, and alerting at scale
- Experience building internal platforms, developer tools, or automation that scales
- Git/version control and CI/CD pipeline experience
- Infrastructure as code and API design experience
- Track record eliminating toil through intelligent automation
- Production ownership experience (on-call, incident response, observability)
- Systems thinking mindset—understanding how components interact at scale
- Eager to dig into problems and bring proposed solutions to group discussion
- Open to feedback and able to creatively adapt multiple ideas into solutions
- Strong technical writing including high and low-level diagramming techniques
- Analytical skills and careful attention to detail
- Availability for rotational on-call duties outside of standard business hours may be required
Why This Role Is Different
You're a key player transforming a team: You will create and develop key relationships with our engineering teams. You will drive a roadmap to enable and govern solid patterns for observability and resilience.
AI at the Forefront: Work with cutting-edge LLM technology to solve real production and observability problems.
Spec-Driven Operations: For the first time, we can specify how the operations harness is to be used while software is being developed. Help define that standard.
Leadership Exposure: Grow into technical acumen, work with leadership across all levels, and shape observability and reliability strategy.
We are in the greatest time ever to be a serious, battle-hardened operator. If you are excited about creating, enabling, and championing world-class observability and response, bring it!
Drug Testing:
To be employed in this role, you'll need to clear a pre-employment drug test. Cox Automotive does not currently administer a pre-employment drug test for marijuana for this position. However, we are a drug-free workplace, so the possession, use or being under the influence of drugs illegal under federal or state law during work hours, on company property and/or in company vehicles is prohibited.
Compensation:
Compensation includes a base salary in the range of $92,300.00 - $153,900.00. The base salary may vary within the anticipated base pay range based on factors such as the ultimate location of the position and the selected candidate’s knowledge, skills, and abilities. Position may be eligible for additional compensation that may include an incentive program.
Benefits:
The Company offers eligible employees the flexibility to take as much vacation with pay as they deem consistent with their duties, the company’s needs, and its obligations; seven paid holidays throughout the calendar year; and up to 160 hours of paid wellness annually for their own wellness or that of family members. Employees are also eligible for additional paid time off in the form of bereavement leave, time off to vote, jury duty leave, volunteer time off, military leave, and parental leave.
EOE, including disability/vets


























