SR. DEVOPS ENGINEER
Make your next big career move by applying as KMC Solutions’ next SR. DEVOPS ENGINEER
We are looking for a DevOps Engineer to own and evolve the infrastructure that powers Anervea's AI platform and supporting services. You will be responsible for designing, deploying, and maintaining a reliable, secure, and cost-efficient cloud environment on AWS, with Docker-based workloads at its core.
This is a hands-on role for someone who enjoys the full lifecycle of infrastructure work — from provisioning and automation to monitoring, incident response, and continuous improvement. You will partner closely with engineering, AI/ML, and product teams to ship reliably and scale gracefully.
The main responsibilities of a SR. DEVOPS ENGINEER include:
Tech Stack You’ll Support
- You will be deploying, scaling, and maintaining infrastructure for the following stack:
- Mobile applications built with Flutter.
- Web frontend built with React.
- Backend services built with Python (FastAPI).
- AI/ML workloads including LLM inference, RAG pipelines, and supporting services.
- Cloud platform: AWS (primary).
- Containerization with Docker; orchestration via ECS or EKS.
- Databases and storage: PostgreSQL/RDS, S3, vector stores, and caching layers (Redis).
Key Responsibilities
AWS Infrastructure
- Design, provision, and maintain AWS infrastructure across services including EC2, ECS/EKS, S3, RDS, VPC, IAM, CloudFront, Route 53, ELB/ALB, Lambda, and CloudWatch.
- Implement Infrastructure as Code (IaC) using Terraform, CloudFormation, or AWS CDK to ensure repeatable and version-controlled environments.
- Manage networking (VPCs, subnets, security groups, NAT gateways, VPN/peering) with a strong focus on security and least-privilege access.
- Optimize AWS spend through right-sizing, reserved instances/savings plans, autoscaling, and continuous cost monitoring
Docker & Containerization
- Build, optimize, and maintain Docker images for backend services, AI/ML workloads, and supporting tooling.
- Manage container orchestration on Amazon ECS (Fargate/EC2) or EKS (Kubernetes), including service definitions, task scaling, and rolling deployments.
- Maintain private container registries (ECR) with tagging, scanning, and lifecycle policies.
- Troubleshoot container-level issues across networking, storage, and runtime performance.
CI/CD & Automation - Build and maintain CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, or AWS CodePipeline / CodeBuild.
- Automate testing, container builds, image scanning, and zero-downtime deployments to staging and production environments.
- Standardize deployment workflows across services so engineers can ship safely and quickly.
Monitoring, Reliability & Incident Response
- Implement and maintain observability across the stack using CloudWatch, Prometheus, Grafana, ELK/OpenSearch, Datadog, or similar.
- Define and track SLOs/SLAs, set up actionable alerting, and reduce noise in on-call workflows.
- Lead incident response, conduct root-cause analyses, and drive postmortems with clear follow-up actions.
- Plan and test backup, disaster recovery, and high-availability strategies.
Security & Compliance
- Apply security best practices across IAM, secrets management (AWS Secrets Manager / Parameter Store), encryption at rest and in transit, and network security.
- Manage SSL/TLS certificates, WAF rules, and DDoS protection.
- Support compliance, audit, and data-protection requirements relevant to healthcare/AI workloads (e.g., data residency, access logging)
Generative AI Infrastructure
- Support and maintain infrastructure for LLM and generative AI workloads on AWS, including Amazon Bedrock, SageMaker, and self-hosted model endpoints.
- Help manage GPU-based compute (EC2 P/G instances, Inferentia) for model inference and training workloads, with a focus on cost and throughput optimization.
- Set up and operate vector databases (OpenSearch, pgvector, Pinecone) and supporting RAG infrastructure.
- Build secure, observable pipelines for model deployment, versioning, and rollback.
- Monitor token usage, inference latency, and model-serving costs; implement guardrails and rate limiting where needed.
- Manage API gateways and authentication for AI endpoints exposed to internal and external consumers.
Collaboration - Work closely with backend, frontend, and AI/ML engineers to understand workload requirements and provide reliable infrastructure.
- Document infrastructure, runbooks, and operational procedures so the team can operate confidently.
- Mentor engineers on DevOps practices, deployment hygiene, and cloud-cost awareness.
Application and Software Development
Applying takes about a minute
Know someone for this?
Refer them in a few clicks and track their progress from your referrals dashboard.