UJET
5 months ago
Senior Site Reliability Engineer
Sign up to save this job, get alerts, and apply with an optimized CV.
Company information
- Company
- UJET
- Location
- Austin, TX, US United States
- Posted
- 5 months ago
Job description
About Us
UJET leads the way in AI-powered contact center innovation, delivering a future-proof, cloud platform that redefines the customer experience with cutting-edge AI, true multimodality, and a mobile-first approach. We infuse AI across every aspect of your customer journey and contact center operations, to drive automation and efficiency. UJET's AI solutions empower agents, optimize customer journeys, and transform contact center operations for elevated experiences and actionable insights. Built on a cloud-native architecture with a unique CRM-first approach, UJET ensures unmatched security, scalability, and prioritized data insights (without storing PII). Designed for effortless use, UJET partners with businesses to deliver exceptional interactions, smarter decision-making, and accelerated growth in the AI-driven world.
Learn more at www.ujet.cx.
Position Overview
Weâre looking for a Senior Site Reliability Engineer to help build and scale a high-impact SRE function. Youâll be a technical leader on a team responsible for improving system reliability, reducing operational toil, and establishing best practices across engineering.bIn this position, youâll design how reliability works in UJET, influence engineering decisions, and build the tooling and processes that make production safer and more predictable.
Responsibilities
- Lead efforts to improve system reliability, scalability, and performance across critical services
- Define and implement SLIs/SLOs and error budgets, and use them to guide engineering priorities
- Design and develop observability systems (metrics, logging, tracing, alerting) that produce actionable alerts and data with minimal noise
- Lead complex incident response, acting as incident commander when needed
- Conduct postmortems focused on systemic causes rather than individual fault, and ensure corrective actions from those reviews are completed.
- Identify and eliminate toil through automation, tooling, and improved workflows
- Partner with product and platform teams on architecture decisions, production readiness, and deployment strategies
- Conduct root cause analysis and implement corrective actions to prevent future incidents
- Develop and maintain a deep understanding of UJET's technology stack and infrastructure
- Collaborate with cross-functional teams to drive reliability and efficiency initiatives
- Stay up-to-date with industry trends, best practices, and emerging technologies to continuously improve UJET's reliability and efficiency
Required skills
- engineering
- security
- incident response
- reliability
- ai
- automation
- cloud-native
- corrective actions
- sre
- infrastructure
- ai-powered
- workflows
- site reliability engineer
- observability
- root cause analysis
- performance
- scalability
- tooling
- contact center
- slis
- alerting
- logging
- metrics
- data insights
- mobile-first
- architecture decisions
- cloud platform
- slos
- technology stack
- production readiness
- tracing
- deployment strategies
- pii
- error budgets
- postmortems
- multimodality
- crm-first
Interested in this position?
Create your free account and tailor your CV to match this job.