Lead SRE observability - SRE Obs G10
US Citizen or GC Holder Required - FedRAMP Requirement
REMOTE
5 days per week/ 8 hours per day
Technology Stack :
Unix, Splunk Enterprise, Splunk Cloud, Elasticsearch, ELK, Kibana, Prometheus, Grafana, Grafana Tempo, OpenTelemetry, Distributed Tracing, Kafka, Terraform, Kubernetes, Docker, Linux, Python, Go, Ruby, Bash, AWS, Ansible, Consul.
Experience supporting FedRAMP or regulated environments.
About the Role
Join our Observability team responsible for designing, building, and operating enterprise platforms for logging, metrics, tracing, and alerting across large-scale cloud infrastructure. You'll lead initiatives that improve reliability, scalability, and operational excellence.
Key Responsibilities
• Design, deploy, and operate enterprise observability platforms.
• Build and maintain Splunk Enterprise/Splunk Cloud infrastructure including Indexers, Search Head Clusters, Heavy Forwarders, and Deployment Servers.
• Deploy and operate large-scale Elasticsearch clusters for log analytics and search.
• Design, deploy, and support distributed tracing platforms using Grafana Tempo and OpenTelemetry.
• Build and maintain end-to-end tracing pipelines, instrumentation standards, and trace retention strategies.
• Scale Prometheus, Grafana, Kafka, Tempo, and OpenTelemetry-based monitoring solutions.
• Develop dashboards, alerts, analytics, and trace visualizations using Splunk SPL, Grafana, Kibana, and Tempo.
• Automate infrastructure using Terraform and configuration management tools.
Required Qualifications
• 7+ years in Site Reliability Engineering, Platform Engineering, or DevOps.
• Hands-on experience administering Splunk Enterprise or Splunk Cloud.
• Strong knowledge of Splunk SPL.
• Experience with Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo, distributed tracing, OpenTelemetry, and Kafka.
• Experience implementing metrics, logs, and traces as part of a modern observability strategy.
• Experience with Terraform and Infrastructure as Code.
• Programming experience in Python, Go, Ruby, or Bash.
Preferred Qualifications
• Splunk certification.
• Experience with Kubernetes, AWS/Azure/GCP, Ansible, Consul, CI/CD pipelines, and service mesh technologies.
• Experience supporting FedRAMP or regulated environments.
REMOTE
5 days per week/ 8 hours per day
Technology Stack :
Unix, Splunk Enterprise, Splunk Cloud, Elasticsearch, ELK, Kibana, Prometheus, Grafana, Grafana Tempo, OpenTelemetry, Distributed Tracing, Kafka, Terraform, Kubernetes, Docker, Linux, Python, Go, Ruby, Bash, AWS, Ansible, Consul.
Experience supporting FedRAMP or regulated environments.
About the Role
Join our Observability team responsible for designing, building, and operating enterprise platforms for logging, metrics, tracing, and alerting across large-scale cloud infrastructure. You'll lead initiatives that improve reliability, scalability, and operational excellence.
Key Responsibilities
• Design, deploy, and operate enterprise observability platforms.
• Build and maintain Splunk Enterprise/Splunk Cloud infrastructure including Indexers, Search Head Clusters, Heavy Forwarders, and Deployment Servers.
• Deploy and operate large-scale Elasticsearch clusters for log analytics and search.
• Design, deploy, and support distributed tracing platforms using Grafana Tempo and OpenTelemetry.
• Build and maintain end-to-end tracing pipelines, instrumentation standards, and trace retention strategies.
• Scale Prometheus, Grafana, Kafka, Tempo, and OpenTelemetry-based monitoring solutions.
• Develop dashboards, alerts, analytics, and trace visualizations using Splunk SPL, Grafana, Kibana, and Tempo.
• Automate infrastructure using Terraform and configuration management tools.
Required Qualifications
• 7+ years in Site Reliability Engineering, Platform Engineering, or DevOps.
• Hands-on experience administering Splunk Enterprise or Splunk Cloud.
• Strong knowledge of Splunk SPL.
• Experience with Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo, distributed tracing, OpenTelemetry, and Kafka.
• Experience implementing metrics, logs, and traces as part of a modern observability strategy.
• Experience with Terraform and Infrastructure as Code.
• Programming experience in Python, Go, Ruby, or Bash.
Preferred Qualifications
• Splunk certification.
• Experience with Kubernetes, AWS/Azure/GCP, Ansible, Consul, CI/CD pipelines, and service mesh technologies.
• Experience supporting FedRAMP or regulated environments.
This is a remote position.
Compensation: $60.00 - $65.00 per hour
WELCOME TO IPOLARITY LLC
IPolarity LLC. is a Professional Services firm composed of highly trained professionals with a wide range of experience in various industries. We provide our clients in key industry verticals with focused expertise in IT Solution, Application Development, Application Maintenance, Testing services and Systems Integration. IPolarity LLC integrates expert industry knowledge, process and technology frameworks, strong partnerships, and a reliable work force to provide strategic solutions that generate sustainable results. We are able to leverage our expertise via a flexible delivery model, which includes onsite and offsite resources.
(if you already have a resume on Indeed)
