Login Enter

Serasa Experian

1 year ago

Full-Stack Site Reliability Engineer (SRE)

Sign up free Log in

Sign up to save this job, get alerts, and apply with an optimized CV.

Company information

Company
Serasa Experian
Location
Brasil, Sudeste, Estado de São Paulo, São Carlos Brazil
Posted
1 year ago
View all jobs at Serasa Experian

Job description



Description of work

About the job

We are looking for a Site Reliability Engineer (SRE) Full-Stack to work in a dynamic and highly collaborative environment, with focus on cloud computing (AWS), microservices, Kubernetes, and infrastructure as code. This position is essential for supporting squads in delivering scalable, secure, and resilient products, as well as contributing directly to audits, observability, and good security practices.

What we are looking for in you

  • Experience with distributed architecture, virtualization, and cloud environments AWS.
  • Mastery of Linux and knowledge of Windows Server.
  • Familiarity with Docker and container orchestration via Kubernetes / AWS EKS.
  • Knowledge of messaging tools like Kafka, RabbitMQ, SQS.
  • Familiarity with DevOps practices, CI/CD (Jenkins), and automation with Terraform, Terragrunt, and Ansible.
  • Ability to work with Shell Script and develop in Python for automations and technical solutions.
  • Experience with agile methodologies (Kanban, Scrum), active participation in ceremonies like daily meetings.
  • Knowledge of AWS products like RDS, EC2, EMR, MWA, as well as banks like DocumentDB.
  • Good communication, proactivity, and ability to work as a team.

How will your day be

  • Work directly in a multidisciplinary squad, participating in agile ceremonies, technical discussions, and strategic decisions.
  • Ensure the quality, security, and resilience of the infrastructure products.
  • Communicate and document the infrastructure design clearly and accessibly to the team.

Main responsibilities

  • Define and implement the infrastructure products according to architecture guidelines.
  • Monitor and ensure resilience, performance, and availability of environments.
  • Manage and align SLIs, SLAs, and SLOs.
  • Perform troubleshooting of infrastructure and support development teams in resolving issues.
  • Suggest and implement solutions for monitoring, logging, and automation.
  • Document the infrastructure and track costs and capacity of environments.
  • Participate in POCs and testing new solutions.
  • Implement infrastructure as code (IaC) with Terraform, Terragrunt, or Ansible.
  • SUPPORT REQUESTS AND INTEGRATIONS WITH INFRASTRUCTURE ON-PREMISES.

Differentiators

  • Experience in high-criticality environments and compliance.
  • Knowledge of observability with tools like Prometheus, Grafana, Datadog or similar.
  • Practical experience in correcting vulnerabilities in infrastructure, containers, and serverless applications.
  • Able to analyze security reports and apply corrections efficiently and securely.
  • Knowledge of hardening, patch management, and risk mitigation in distributed environments.
  • Familiarity with scanning and compliance tools integrated into the development cycle and infrastructure as code.

Required skills

Interested in this position?

Create your free account and tailor your CV to match this job.