AI DevOps Engineer
Il y a 1 semaine
Tunis, Tunisie
VegaNext
Temps plein
Gratuit avec email ou Google
Enregistrez cette offre et organisez votre recherche
Créez un compte gratuit pour enregistrer des offres d'emploi, créer des alertes et revenir à cette liste depuis votre tableau de bord.
Gratuit avec email ou Google
En continuant, vous acceptez nos Conditions d’utilisation & Politique de confidentialité.
About VegaNext
At VegaNext, we build and operate production AI systems, including real-time AI agents that handles multi-step conversations end to end (understanding intent, managing state, and confirming outcomes). These are live production systems used by real users on a daily basis.
Role Summary
As an AI DevOps Engineer, you'll work across two connected areas: contributing to the reliability of our AI
agent systems, and helping build and maintain the infrastructure, pipelines, and cloud operations that keep our platforms running.
What You'll Do
AI Agent & Reliability Work
• Work on AI agent behavior: prompts, conversation flows, tool calling
• Build tests that replay realistic conversations and verify outcomes
• Investigate transcripts and logs to understand why interactions fail
• Debug and fix issues across the Python backend and, occasionally, the dashboard Infrastructure & DevOps
• Build and maintain CI/CD pipelines for deployment across services
• Manage cloud infrastructure (provisioning, scaling, cost and security hygiene)
• Write and maintain infrastructure as code (Terraform, CloudFormation, or similar)
• Set up and improve monitoring, logging, and alerting across production systems
• Containerize services and manage orchestration (Docker, and container runtime environments)
• Support incident response: diagnose production issues and help drive fixes
• Continuously look for ways to automate manual operational work What We're Looking For
• Solid Python skills and the ability to navigate an unfamiliar codebase
• Hands-on experience with CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, or similar)
• Working knowledge of cloud infrastructure (AWS, GCP, or Azure)
• Comfortable with containerization (Docker) and basic orchestration concepts
• Comfortable working with SQL and REST APIs
• Familiarity with Git and software testing
• Strong debugging and problem-solving skills across both application and infrastructure layers
• Curiosity and attention to detail: you verify what a system actually does rather than assuming it works as expected Nice to Have
• Infrastructure as code experience (Terraform, Pulumi, CloudFormation)
• Experience with LLM APIs: prompting, function/tool calling, structured outputs
• Voice or real-time audio: LiveKit, Twilio, STT, TTS, WebRTC
• Kubernetes or other container orchestration
• Monitoring/observability tooling (Prometheus, Grafana, Datadog, or similar)
• Experience evaluating or testing LLM/agent behavior
What You'll Do
AI Agent & Reliability Work
• Work on AI agent behavior: prompts, conversation flows, tool calling
• Build tests that replay realistic conversations and verify outcomes
• Investigate transcripts and logs to understand why interactions fail
• Debug and fix issues across the Python backend and, occasionally, the dashboard Infrastructure & DevOps
• Build and maintain CI/CD pipelines for deployment across services
• Manage cloud infrastructure (provisioning, scaling, cost and security hygiene)
• Write and maintain infrastructure as code (Terraform, CloudFormation, or similar)
• Set up and improve monitoring, logging, and alerting across production systems
• Containerize services and manage orchestration (Docker, and container runtime environments)
• Support incident response: diagnose production issues and help drive fixes
• Continuously look for ways to automate manual operational work What We're Looking For
• Solid Python skills and the ability to navigate an unfamiliar codebase
• Hands-on experience with CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, or similar)
• Working knowledge of cloud infrastructure (AWS, GCP, or Azure)
• Comfortable with containerization (Docker) and basic orchestration concepts
• Comfortable working with SQL and REST APIs
• Familiarity with Git and software testing
• Strong debugging and problem-solving skills across both application and infrastructure layers
• Curiosity and attention to detail: you verify what a system actually does rather than assuming it works as expected Nice to Have
• Infrastructure as code experience (Terraform, Pulumi, CloudFormation)
• Experience with LLM APIs: prompting, function/tool calling, structured outputs
• Voice or real-time audio: LiveKit, Twilio, STT, TTS, WebRTC
• Kubernetes or other container orchestration
• Monitoring/observability tooling (Prometheus, Grafana, Datadog, or similar)
• Experience evaluating or testing LLM/agent behavior