InfrastructureAWSPlatformCI/CDKubernetesTerraformAnsibleCloudFormationJenkinsGitlab CI
📋 Job Overview
At Rehire, we are partnering with a US-based data engineering and cloud technologies company to find a Junior SRE Engineer to join its Reliability Engineering team. You will be embedded with the SRE function of a financial services client, supporting a regulated consumer-lending platform on AWS with hundreds of microservices and event-driven pipelines, where reliability directly impacts customer trust and compliance.
🏢 About [Company]
Rehire is a US-based data engineering and cloud technologies company specializing in innovative solutions for financial services clients.
🎯 The Role
You will be embedded with the SRE function of a financial services client, supporting a regulated consumer-lending platform on AWS with hundreds of microservices and event-driven pipelines, where reliability directly impacts customer trust and compliance. This is a unique opportunity to grow in an environment where AI is already part of day-to-day operations, including an AI SRE co-pilot, purpose-built AI agents for incident triage and monitoring, and a formal AI governance program. Working under the guidance of senior engineers and architects, you will help operate, improve, and learn from this AI-driven reliability program.
✅ Key Responsibilities
Support incident response as a shadow or secondary responder, using AI-driven detection, correlation, and root-cause analysis tools.
Help build incident timelines from metrics, logs, traces, and deploy history, and document postmortems and corrective actions through to closure.
Help operate and monitor AI SRE sub-agents (incident summarization, monitor-gap detection, usage attribution), flagging anomalies for senior review.
Build and maintain Datadog monitors, dashboards, and SLO definitions as code using Terraform, and support SLI/SLO and error-budget reviews for critical customer journeys.
Help reduce alert noise with AI/ML-assisted detection (anomaly, outlier, and forecast monitors, dynamic thresholds) and run recurring monitor-hygiene reviews.
Support the day-to-day reliability of AWS workloads (ECS/Fargate, EKS, Lambda, RDS/Aurora, ALB, SQS/SNS, Step Functions).
Identify capacity, saturation, and cloud cost anomalies, and help attribute spend and telemetry volume to owning teams and services.
Write automation in Python and Bash against platform APIs (Datadog, AWS, GitHub, PagerDuty, Jira) and contribute Terraform modules through pull requests.
Help integrate reliability controls and AI-assisted checks into CI/CD pipelines, and create runbooks progressively automated toward self-healing.
Participate in architecture, reliability, and AI-risk reviews, learning how compliance frameworks (PCI-DSS, SOC 2, SOX, GLBA) apply in a regulated environment.
📌 Required Qualifications
Bachelor's degree in Computer Science, Engineering, or a related field.
1–3 years of experience in SRE, DevOps, cloud infrastructure, platform, or production-support engineering.
Advanced English (oral and written). REQUIRED
Hands-on exposure to AWS (or an equivalent hyperscaler) across compute, networking, storage, and managed database services.
Exposure to at least one observability platform (Datadog preferred; Grafana/Prometheus, New Relic, CloudWatch, or ELK/OpenSearch also relevant).
Foundational understanding of SLI, SLO, and error-budget concepts.
Foundational knowledge of containers and orchestration (Docker, ECS, or Kubernetes) and serverless execution models.
Beginner-to-intermediate experience with Infrastructure as Code (Terraform preferred; Ansible or CloudFormation acceptable).
Scripting experience in Python, Bash, or similar, including consuming REST APIs and parsing JSON.
Basic Linux troubleshooting and networking fundamentals (DNS, TLS, load balancing, timeouts, and retries).
Comfort with Git, pull-request workflows, and CI/CD tools (GitHub Actions, Jenkins, GitLab CI, ArgoCD, or similar).
Familiarity with incident management and on-call concepts (severity models, escalation policies, PagerDuty or Opsgenie).
⭐ Desirable Experience
Experience in product engineering services, enterprise software, or fintech is a plus.
Awareness of compliance frameworks (PCI-DSS, SOC 2, SOX, GLBA) is a plus.
🎁 Benefits
Competitive Salary Paid in USD. Dynamic and collaborative work environment. Hands-on learning in AI-driven SRE practices and opportunities for career advancement.
🛂 Visa & Eligibility
Work Schedule: US shift presential at Guadalajara, Mexico. Work Modality: Full-time contractor basis.