Urgently hiring Use left and right arrow keys to navigate
Based on similar jobs in your market
Estimated Pay info$48 per hour
Hours Full-time
Location Herndon, VA
Herndon, Virginia open_in_new

About this job

Job Description

Job Description

Contract Details

  • Work Mode: 100% Remote (US-based)
  • Location: Herndon, VA
  • Schedule: 40 hours/week
  • Duration: 08/17/2026 08/16/2027
  • Type: Contract with potential to convert to full-time after ~12 months (not guaranteed)

About the Opportunity

Seeking a Site Reliability Engineer to ensure availability, performance, scalability, and security for mission-critical, cloud-hosted search and analytics services built on OpenSearch. You will focus on reliability engineering, operations, automation, and continuous improvement for distributed platforms, working within a diverse, globally distributed team.

Key Responsibilities

  • Provision, build, deploy, monitor, operate, and support cloud services in a global team environment.
  • Architect, build, deploy, and maintain high performance OpenSearch clusters and platforms from the ground up.
  • Optimize OpenSearch for high availability, resiliency, scalability, security, and performance.
  • Monitor and troubleshoot cluster health, node performance, indexing throughput, search latency, shard allocation/replication, and storage utilization.
  • Analyze and resolve operational issues across infrastructure, platform, and application layers; lead incident response, RCA, and remediation.
  • Maintain integrity and security of servers, systems, and OpenSearch platform infrastructure.
  • Support lifecycle activities: installation, configuration, upgrades/patching, backup/restore, and disaster recovery.
  • Develop and maintain monitoring policies, alerting standards, runbooks, and support procedures.
  • Automate testing, deployment, scaling, recovery, and operational workflows for OpenSearch and related cloud services.
  • Plan capacity for compute, memory, storage, and network; partner with engineering to enhance reliability and operational readiness.
  • Support log ingestion, index management, lifecycle/retention, and search performance tuning.
  • Participate in an on-call rotation; support occasional weekend/after-hours needs.

Required Qualifications

  • US citizenship required; dual citizenship not permitted.
  • 8 years of experience in SRE/DevOps/cloud operations with distributed systems.
  • Proven, hands-on experience designing, building, deploying, operating, and optimizing OpenSearch clusters from scratch in production.
  • Expert-level Kubernetes experience (operations, troubleshooting, management, configuration of complex services).
  • Deep OpenSearch administration: cluster architecture, performance tuning, scaling, upgrades, and troubleshooting; index/shard/replica strategy; sizing; snapshot/restore; backup/DR.
  • Strong Linux expertise (SUSE and Ubuntu).
  • Expertise with Git and Concourse (pipeline setup, management, troubleshooting).
  • Experience with Kafka and Zookeeper; strong automation for testing, deployment, scalability, and cloud service management.
  • Experience building/implementing/supporting cloud monitoring and observability; solid knowledge of cloud computing, infrastructure operations, databases, web services, networking, virtualization, and internet protocols.
  • Security fundamentals for SaaS multi-tenant application systems; excellent communication and prioritization skills; ability to multitask.

Preferred Qualifications

  • AWS experience (e.g., Route 53, EC2, S3, CloudWatch, DynamoDB, RDS, IAM, ACM, KMS, VPC); experience deploying/operating OpenSearch in AWS.
  • Experience with Cloud Foundry environments.
  • Experience with Jenkins, Chef, and/or Terraform.
  • Experience with Prometheus and Grafana.
  • Background with log ingestion pipelines, index lifecycle management, retention strategies, and search platform security controls.
  • Familiarity with capacity forecasting, performance benchmarking, and resilience testing for distributed search platforms.

Work Environment

  • Collaborative, globally distributed team with cross-training opportunities.
  • Participation in an on-call rotation and occasional after-hours/weekend support.
#ZR

Nearby locations

Posting ID: 1286349665 Posted: 2026-08-08 Job Title: Site Reliability Engineer