AI DevOps Engineer

Il y a 1 semaine

Tunis, Tunisie VegaNext Temps plein
About VegaNext At VegaNext, we build and operate production AI systems, including real-time AI agents that handles multi-step conversations end to end (understanding intent, managing state, and confirming outcomes). These are live production systems used by real users on a daily basis. Role Summary As an AI DevOps Engineer, you'll work across two connected areas: contributing to the reliability of our AI agent systems, and helping build and maintain the infrastructure, pipelines, and cloud operations that keep our platforms running.

What You'll Do
AI Agent & Reliability Work
• Work on AI agent behavior: prompts, conversation flows, tool calling
• Build tests that replay realistic conversations and verify outcomes
• Investigate transcripts and logs to understand why interactions fail
• Debug and fix issues across the Python backend and, occasionally, the dashboard Infrastructure & DevOps
• Build and maintain CI/CD pipelines for deployment across services
• Manage cloud infrastructure (provisioning, scaling, cost and security hygiene)
• Write and maintain infrastructure as code (Terraform, CloudFormation, or similar)
• Set up and improve monitoring, logging, and alerting across production systems
• Containerize services and manage orchestration (Docker, and container runtime environments)
• Support incident response: diagnose production issues and help drive fixes
• Continuously look for ways to automate manual operational work What We're Looking For
• Solid Python skills and the ability to navigate an unfamiliar codebase
• Hands-on experience with CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, or similar)
• Working knowledge of cloud infrastructure (AWS, GCP, or Azure)
• Comfortable with containerization (Docker) and basic orchestration concepts
• Comfortable working with SQL and REST APIs
• Familiarity with Git and software testing
• Strong debugging and problem-solving skills across both application and infrastructure layers
• Curiosity and attention to detail: you verify what a system actually does rather than assuming it works as expected Nice to Have
• Infrastructure as code experience (Terraform, Pulumi, CloudFormation)
• Experience with LLM APIs: prompting, function/tool calling, structured outputs
• Voice or real-time audio: LiveKit, Twilio, STT, TTS, WebRTC
• Kubernetes or other container orchestration
• Monitoring/observability tooling (Prometheus, Grafana, Datadog, or similar)
• Experience evaluating or testing LLM/agent behavior