KMC Careers Logo

SR. DEVOPS ENGINEER

Application and Software DevelopmentPosted Jul 06

Make your next big career move by applying as KMC Solutions’ next SR. DEVOPS ENGINEER

We are looking for a DevOps Engineer to own and evolve the infrastructure that powers Anervea's AI platform and supporting services. You will be responsible for designing, deploying, and maintaining a reliable, secure, and cost-efficient cloud environment on AWS, with Docker-based workloads at its core.

This is a hands-on role for someone who enjoys the full lifecycle of infrastructure work — from provisioning and automation to monitoring, incident response, and continuous improvement. You will partner closely with engineering, AI/ML, and product teams to ship reliably and scale gracefully.

The main responsibilities of a SR. DEVOPS ENGINEER include:

Tech Stack You’ll Support

  • You will be deploying, scaling, and maintaining infrastructure for the following stack:
  • Mobile applications built with Flutter.
  • Web frontend built with React.
  • Backend services built with Python (FastAPI).
  • AI/ML workloads including LLM inference, RAG pipelines, and supporting services.
  • Cloud platform: AWS (primary).
  • Containerization with Docker; orchestration via ECS or EKS.
  • Databases and storage: PostgreSQL/RDS, S3, vector stores, and caching layers (Redis).

Key Responsibilities

AWS Infrastructure

  • Design, provision, and maintain AWS infrastructure across services including EC2, ECS/EKS, S3, RDS, VPC, IAM, CloudFront, Route 53, ELB/ALB, Lambda, and CloudWatch.
  • Implement Infrastructure as Code (IaC) using Terraform, CloudFormation, or AWS CDK to ensure repeatable and version-controlled environments.
  • Manage networking (VPCs, subnets, security groups, NAT gateways, VPN/peering) with a strong focus on security and least-privilege access.
  • Optimize AWS spend through right-sizing, reserved instances/savings plans, autoscaling, and continuous cost monitoring

Docker & Containerization

  • Build, optimize, and maintain Docker images for backend services, AI/ML workloads, and supporting tooling.
  • Manage container orchestration on Amazon ECS (Fargate/EC2) or EKS (Kubernetes), including service definitions, task scaling, and rolling deployments.
  • Maintain private container registries (ECR) with tagging, scanning, and lifecycle policies.
  • Troubleshoot container-level issues across networking, storage, and runtime performance.
    CI/CD & Automation
  • Build and maintain CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, or AWS CodePipeline / CodeBuild.
  • Automate testing, container builds, image scanning, and zero-downtime deployments to staging and production environments.
  • Standardize deployment workflows across services so engineers can ship safely and quickly.

Monitoring, Reliability & Incident Response

  • Implement and maintain observability across the stack using CloudWatch, Prometheus, Grafana, ELK/OpenSearch, Datadog, or similar.
  • Define and track SLOs/SLAs, set up actionable alerting, and reduce noise in on-call workflows.
  • Lead incident response, conduct root-cause analyses, and drive postmortems with clear follow-up actions.
  • Plan and test backup, disaster recovery, and high-availability strategies.


Security & Compliance

  • Apply security best practices across IAM, secrets management (AWS Secrets Manager / Parameter Store), encryption at rest and in transit, and network security.
  • Manage SSL/TLS certificates, WAF rules, and DDoS protection.
  • Support compliance, audit, and data-protection requirements relevant to healthcare/AI workloads (e.g., data residency, access logging)

Generative AI Infrastructure

  • Support and maintain infrastructure for LLM and generative AI workloads on AWS, including Amazon Bedrock, SageMaker, and self-hosted model endpoints.
  • Help manage GPU-based compute (EC2 P/G instances, Inferentia) for model inference and training workloads, with a focus on cost and throughput optimization.
  • Set up and operate vector databases (OpenSearch, pgvector, Pinecone) and supporting RAG infrastructure.
  • Build secure, observable pipelines for model deployment, versioning, and rollback.
  • Monitor token usage, inference latency, and model-serving costs; implement guardrails and rate limiting where needed.
  • Manage API gateways and authentication for AI endpoints exposed to internal and external consumers.
    Collaboration
  • Work closely with backend, frontend, and AI/ML engineers to understand workload requirements and provide reliable infrastructure.
  • Document infrastructure, runbooks, and operational procedures so the team can operate confidently.
  • Mentor engineers on DevOps practices, deployment hygiene, and cloud-cost awareness.

Application and Software Development

Applying takes about a minute

Know someone for this?

Refer them in a few clicks and track their progress from your referrals dashboard.