Serasa Experian
1 year ago
Full-Stack Site Reliability Engineer (SRE)
Sign up to save this job, get alerts, and apply with an optimized CV.
Company information
- Company
- Serasa Experian
- Location
- Brasil, Sudeste, Estado de São Paulo, São Carlos Brazil
- Posted
- 1 year ago
Job description
Description of work
About the job
We are looking for a Site Reliability Engineer (SRE) Full-Stack to work in a dynamic and highly collaborative environment, with focus on cloud computing (AWS), microservices, Kubernetes, and infrastructure as code. This position is essential for supporting squads in delivering scalable, secure, and resilient products, as well as contributing directly to audits, observability, and good security practices.
What we are looking for in you
- Experience with distributed architecture, virtualization, and cloud environments AWS.
- Mastery of Linux and knowledge of Windows Server.
- Familiarity with Docker and container orchestration via Kubernetes / AWS EKS.
- Knowledge of messaging tools like Kafka, RabbitMQ, SQS.
- Familiarity with DevOps practices, CI/CD (Jenkins), and automation with Terraform, Terragrunt, and Ansible.
- Ability to work with Shell Script and develop in Python for automations and technical solutions.
- Experience with agile methodologies (Kanban, Scrum), active participation in ceremonies like daily meetings.
- Knowledge of AWS products like RDS, EC2, EMR, MWA, as well as banks like DocumentDB.
- Good communication, proactivity, and ability to work as a team.
How will your day be
- Work directly in a multidisciplinary squad, participating in agile ceremonies, technical discussions, and strategic decisions.
- Ensure the quality, security, and resilience of the infrastructure products.
- Communicate and document the infrastructure design clearly and accessibly to the team.
Main responsibilities
- Define and implement the infrastructure products according to architecture guidelines.
- Monitor and ensure resilience, performance, and availability of environments.
- Manage and align SLIs, SLAs, and SLOs.
- Perform troubleshooting of infrastructure and support development teams in resolving issues.
- Suggest and implement solutions for monitoring, logging, and automation.
- Document the infrastructure and track costs and capacity of environments.
- Participate in POCs and testing new solutions.
- Implement infrastructure as code (IaC) with Terraform, Terragrunt, or Ansible.
- SUPPORT REQUESTS AND INTEGRATIONS WITH INFRASTRUCTURE ON-PREMISES.
Differentiators
- Experience in high-criticality environments and compliance.
- Knowledge of observability with tools like Prometheus, Grafana, Datadog or similar.
- Practical experience in correcting vulnerabilities in infrastructure, containers, and serverless applications.
- Able to analyze security reports and apply corrections efficiently and securely.
- Knowledge of hardening, patch management, and risk mitigation in distributed environments.
- Familiarity with scanning and compliance tools integrated into the development cycle and infrastructure as code.
Required skills
Interested in this position?
Create your free account and tailor your CV to match this job.