Archived listing. This role was posted over 30 days ago. We keep it for reference, but the employer may have already filled it. See today's verified listings.

Senior Elasticsearch Engineer

Chess.com Remote

Posted 20 Jul 2026
Last seen 04 Aug 2026
Location Remote
Lifecycle mature
Grade D

Engineering

About YouChess.com is the world's largest chess platform with 235M+ members and ~20 million daily games. Our Elasticsearch and OpenSearch infrastructure underpins search, user activity, analytics, logging, and operational intelligence at massive scale, hundreds of terabytes across a dozen production clusters running on bare-metal Kubernetes.We're looking for a Senior Elasticsearch Engineer who can own the full lifecycle of our search and analytics data platform: capacity planning, cluster architecture, performance tuning, incident response, migration strategy, and operational excellence. You'll be the single point of deep expertise across all Elasticsearch and OpenSearch clusters at Chess.com.This is not a monitoring-from-dashboards role. You'll be hands-on with cluster internals, write ILM/ISM policies, push infrastructure changes through GitOps, and make real-time decisions about replica allocation when a cluster goes red.What you'll doIncident Response & ReliabilityShard allocation strategy for write-heavy data streams at high throughput (millions of documents per minute)Disk watermark management, retention policy tuning, and rollover orchestration for high-volume indicesPerformance optimization and I/O tuning on bare-metal nodesWrite queue analysis, thread pool diagnostics, and shard rebalancing under loadCapacity planning and growth forecasting across clustersIncident Response & ReliabilityOn-call ownership for Elasticsearch-related incidents: cluster health degradation, node loss, disk pressure, shard imbalance, and write rejection cascadesReal-time cluster triage and cross-team coordination during production incidentsPost-mortem authoring and systemic reliability improvementsSnapshot and disaster recovery management across clustersMigration & StrategyElasticsearch-to-OpenSearch migration analysis and execution, including compatibility evaluation across ILM/ISM, security models, and plugin ecosystemsVersion upgrade planning and rolling restart orchestration with zero-downtime requirementsEnd-to-end new cluster provisioning and onboardingCross-Team EnablementAdvise engineering teams on index design, mapping strategy, retention policies, and query optimizationManage Kibana and OpenSearch Dashboards access and configuration for internal consumersDefine and maintain workload priority tiers across clustersPreferred Skills7+ years operating Elasticsearch at scale (multi-TB clusters, dozens of nodes, high write throughput)Deep understanding of Elasticsearch internals: segment merging, translog, shard allocation, and cluster state managementProduction experience with ECK (Elastic Cloud on Kubernetes) or equivalent operator-based deploymentsProficiency with Kubernetes operations for stateful workloads (StatefulSets, persistent storage, resource management)Hands-on Linux systems administration with a focus on storage and I/O performanceExperience managing both Elasticsearch and OpenSearch in production, including an informed opinion on their respective trade-offsIncident command experience: ability to diagnose and mitigate cluster emergencies under pressure while communicating clearly to stakeholdersGit-based infrastructure management (GitOps): Helm charts, ArgoCD/Flux, infrastructure-as-code for cluster configurationFluency with the Elastic stack APIs: cluster administration, index templates, data streams, ILM policies, snapshot/restoreBonus ExperienceOpenSearch ISM policies and security plugin (fine-grained access control)GCS or S3 snapshot repository configuration and cross-cluster replicationGrafana + Prometheus monitoring for Elasticsearch metricsKibana Discover, Dev Tools, and data view management at scaleJava internals relevant to Elasticsearch JVM tuning (heap sizing, GC tuning, circuit breakers)Vault integration for secrets management in Kubernetes-deployed search clustersFluentd/Fluent Bit log pipeline configuration feeding OpenSearchHardware selection experience for search-optimized server configurationsPython or scripting for operational analysis and automationWhat Makes This Role SpecialFull autonomy. You are the Elasticsearch authority. You make the architecture calls, set the priorities, and own the outcomes.Real scale. Hundreds of terabytes of data, billions of documents, millions of daily queries. The problems here don't exist at smaller companies.Bare metal. No managed Elastic Cloud. You're operating directly on the hardware. This is hands-on engineering.Strategic impact. Your decisions on ES vs. OpenSearch migration, cluster topology, and capacity planning directly affect product capabilities and infrastructure costs.Small team, big trust. Chess.com runs lean. You won't be buried in process or approvals. Ship changes, fix problems, improve systems.About the OpportunityThis is a full-time opportunityWe are 100% remote (work from anywhere!)---You can learn more about us here:https://www.chess.com/article/view/how-chess-com-virtual-team-works-togetherhttps://www.chess.com/about
About UsChess.com is one of the largest gaming sites in the world and the #1 platform for playing, learning, and enjoying chess.We are a team of 600+ fully remote people in 60+ countries working hard to serve the global chess community. We are here to support 250M+ chess players worldwide with the best possible product, content, and tools to serve the community!We are a tech company. A gaming company. A content company. And we do it all with passion and commitment to the game. Above all we prize our mission-driven, flat, life-celebrating, no-corporate culture, and we look forward to meeting you and learning more about what you can bring to the team.

Apply for this role