Observability Platform Engineer
Il y a 2 semaines
Tunis, 11, Tunisie
ODDO BHF
Temps plein
Gratuit avec email ou Google
Enregistrez cette offre et organisez votre recherche
Créez un compte gratuit pour enregistrer des offres d'emploi, créer des alertes et revenir à cette liste depuis votre tableau de bord.
Gratuit avec email ou Google
*
Key Responsibilities
* Observability Platform Administration:
- Administer, maintain, and optimize observability solutions, including:
- Prometheus
- Thanos
- Grafana
- Loki
- Tempo
- OpenTelemetry Collectors
- Ensure the availability, performance, scalability, and reliability of observability services.
- Manage platform configurations, upgrades, and security patching.
- Oversee data retention policies and optimize the storage of metrics, logs, and traces.
- Manage storage capacity planning and data archiving strategies. Developer Support and Enablement:
- Support development teams in implementing application observability best practices.
- Provide technical expertise on metrics, logging, distributed tracing, and alerting.
- Define, promote, and enforce observability standards across the organization.
- Train and mentor teams on monitoring, troubleshooting, and observability best practices.
- Act as a trusted partner for development and operations teams to improve service reliability and visibility. Continuous Improvement:
- Evolve the observability architecture and tooling to meet business and technical requirements.
- Automate recurring administration and configuration tasks.
- Improve the quality and relevance of collected telemetry data and alerts.
- Reduce operational noise and alert fatigue while enhancing the developer experience.
- Identify and implement opportunities to increase platform efficiency, scalability, and resilience. Governance and Documentation:
- Maintain comprehensive technical and operational documentation.
- Establish and manage standards for naming conventions, tagging, monitoring, and dashboard design.
- Ensure compliance with security, operational, and governance requirements.
- Produce regular reports on platform usage, performance, capacity, and overall health. *Qualifications & Experience*
- Bachelor's or Master's degree in Computer Science, Software Engineering, Information Technology, or a related field.
- Minimum 2 years of experience in Observability, DevOps, Site Reliability Engineering (SRE), Platform Engineering, or Cloud Infrastructure environments.
- Hands-on experience with Kubernetes and/or OpenShift platforms.
- Practical knowledge of observability and monitoring solutions, including: Prometheus,Grafana, Loki, Tempo, OpenTelemetry
- Experience supporting production environments and implementing monitoring, logging, tracing, and alerting best practices.
- Familiarity with cloud-native architectures, containerized applications, and automation principles.
- Strong understanding of infrastructure reliability, performance monitoring, and operational excellence.
Key Responsibilities
* Observability Platform Administration:
- Administer, maintain, and optimize observability solutions, including:
- Prometheus
- Thanos
- Grafana
- Loki
- Tempo
- OpenTelemetry Collectors
- Ensure the availability, performance, scalability, and reliability of observability services.
- Manage platform configurations, upgrades, and security patching.
- Oversee data retention policies and optimize the storage of metrics, logs, and traces.
- Manage storage capacity planning and data archiving strategies. Developer Support and Enablement:
- Support development teams in implementing application observability best practices.
- Provide technical expertise on metrics, logging, distributed tracing, and alerting.
- Define, promote, and enforce observability standards across the organization.
- Train and mentor teams on monitoring, troubleshooting, and observability best practices.
- Act as a trusted partner for development and operations teams to improve service reliability and visibility. Continuous Improvement:
- Evolve the observability architecture and tooling to meet business and technical requirements.
- Automate recurring administration and configuration tasks.
- Improve the quality and relevance of collected telemetry data and alerts.
- Reduce operational noise and alert fatigue while enhancing the developer experience.
- Identify and implement opportunities to increase platform efficiency, scalability, and resilience. Governance and Documentation:
- Maintain comprehensive technical and operational documentation.
- Establish and manage standards for naming conventions, tagging, monitoring, and dashboard design.
- Ensure compliance with security, operational, and governance requirements.
- Produce regular reports on platform usage, performance, capacity, and overall health. *Qualifications & Experience*
- Bachelor's or Master's degree in Computer Science, Software Engineering, Information Technology, or a related field.
- Minimum 2 years of experience in Observability, DevOps, Site Reliability Engineering (SRE), Platform Engineering, or Cloud Infrastructure environments.
- Hands-on experience with Kubernetes and/or OpenShift platforms.
- Practical knowledge of observability and monitoring solutions, including: Prometheus,Grafana, Loki, Tempo, OpenTelemetry
- Experience supporting production environments and implementing monitoring, logging, tracing, and alerting best practices.
- Familiarity with cloud-native architectures, containerized applications, and automation principles.
- Strong understanding of infrastructure reliability, performance monitoring, and operational excellence.