Observability Platform Engineer

Il y a 2 semaines

Tunis, 11, Tunisie ODDO BHF Temps plein
*

Key Responsibilities
* Observability Platform Administration:
- Administer, maintain, and optimize observability solutions, including:
- Prometheus
- Thanos
- Grafana
- Loki
- Tempo
- OpenTelemetry Collectors
- Ensure the availability, performance, scalability, and reliability of observability services.
- Manage platform configurations, upgrades, and security patching.
- Oversee data retention policies and optimize the storage of metrics, logs, and traces.
- Manage storage capacity planning and data archiving strategies. Developer Support and Enablement:
- Support development teams in implementing application observability best practices.
- Provide technical expertise on metrics, logging, distributed tracing, and alerting.
- Define, promote, and enforce observability standards across the organization.
- Train and mentor teams on monitoring, troubleshooting, and observability best practices.
- Act as a trusted partner for development and operations teams to improve service reliability and visibility. Continuous Improvement:
- Evolve the observability architecture and tooling to meet business and technical requirements.
- Automate recurring administration and configuration tasks.
- Improve the quality and relevance of collected telemetry data and alerts.
- Reduce operational noise and alert fatigue while enhancing the developer experience.
- Identify and implement opportunities to increase platform efficiency, scalability, and resilience. Governance and Documentation:
- Maintain comprehensive technical and operational documentation.
- Establish and manage standards for naming conventions, tagging, monitoring, and dashboard design.
- Ensure compliance with security, operational, and governance requirements.
- Produce regular reports on platform usage, performance, capacity, and overall health. *Qualifications & Experience*
- Bachelor's or Master's degree in Computer Science, Software Engineering, Information Technology, or a related field.
- Minimum 2 years of experience in Observability, DevOps, Site Reliability Engineering (SRE), Platform Engineering, or Cloud Infrastructure environments.
- Hands-on experience with Kubernetes and/or OpenShift platforms.
- Practical knowledge of observability and monitoring solutions, including: Prometheus,Grafana, Loki, Tempo, OpenTelemetry
- Experience supporting production environments and implementing monitoring, logging, tracing, and alerting best practices.
- Familiarity with cloud-native architectures, containerized applications, and automation principles.
- Strong understanding of infrastructure reliability, performance monitoring, and operational excellence.