Software Engineer, Voice
Salva questo lavoro e mantieni la tua ricerca organizzata
Crea un account gratuito per salvare lavori, creare avvisi e tornare a questa inserzione dalla tua dashboard.
Continuando accetti i nostri Termini & Informativa sulla privacy.
If you are here, it is because you know that we are looking for a Software Engineer, Voice for our Product team.
indigo.ai is the leading platform in Italy for building next-generation AI Agents that transform the way companies communicate with their customers. Since 2016, we’ve been helping enterprises in industries such as finance, insurance, utilities, retail, and e-commerce to evolve their Customer Experience through conversational AI. We don’t just “sell software”: we enable a shift in how organizations interact with people, automating millions of conversations every year, reducing operational costs, improving conversion rates, and creating more personalized, scalable, and compliant customer journeys.
Backed by a recent €10 million investment from Azimut, we are on a mission to take this technology global. This is a unique opportunity to join a well-funded, highly ambitious team and play a direct role in shaping the future of enterprise AI.
To make this happen, we are looking for a Software Engineer, Voice to support our Chief Product Development Officer in making talking to AI on the phone feel human.
What are we looking for?We are looking for a Software Engineer, Voice to own the real-time voice layer of our AI Agents. Voice is where conversational AI is being decided right now, and our voice agents already handle production phone traffic for large companies. The bar is moving fast: we want to make the leap from \"works reliably\" to \"feels human on a real phone line\", and we want one person to own that leap. This is a specialist role with end-to-end ownership: the architecture, the model and provider choices, the latency budget, the way a conversation feels. You'll join our Product Engineering team, reporting to our Chief Product Development Officer, as our first full-time engineer dedicated to voice. And if voice grows the way we believe it will, you'll shape the team that grows around it.
Key Responsibilities:Own the real-time voice pipeline end-to-end. From audio ingress on the telephony edge, through streaming STT and turn-taking, to the agent brain and back out through streaming TTS. Every millisecond in between is yours.
Engineer how fast the agent feels. Semantic end-of-turn detection, preemptive generation on partial transcripts, eager TTS, filler and backchannel strategies that mask tool calls. All measured on real 8kHz phone audio, not in a browser demo.
Make turn-taking human. Barge-in that survives noisy lines. Endpointing policies that know the dialog state, so a caller never gets cut off mid-IBAN. The difference between an IVR and a conversation lives here.
Raise voice quality on the channel that actually ships: the phone. Benchmark and A/B STT and TTS providers on real G.711 calls (Italian first: WER, naturalness, numbers and codes read right), exploit wideband/HD voice where the carrier allows it, and experiment with context-aware TTS and conversational speech models as they mature.
Build the evaluation harness. Turn \"this voice sounds better\" into numbers we trust: per-stage latency budgets, turn-taking metrics, regression suites on recorded calls, quality gates before anything reaches a client.
Keep production boringly reliable. Per-stage observability, live-call incident debugging (dead air, stuck turns, provider hiccups), graceful degradation when a vendor blinks.
Track a weekly-moving ecosystem and turn it into strategy. New STT/TTS/speech-to-speech releases land every month. You decide what we integrate, what we self-host for EU compliance and data residency, and what we skip. And you make provider swaps cheap.
The filter is not your degree, and it's not years-of-experience arithmetic. It's having built it. Tell us about a real-time voice or audio system you designed and shipped: the latency budget, where it broke, and what you changed to make it feel right. That tells us more than any title.
Real-time audio systems, shipped. You've built voice agents, telephony systems, conferencing or live-streaming products that ran in production. You know what it means to move audio over WebSockets/WebRTC/SIP, through codecs (G.711/μ-law, Opus), against a latency budget.
The modern voice AI stack, hands-on. Streaming STT and TTS, VAD and turn detection, voice orchestration frameworks (Pipecat, LiveKit Agents or equivalent), speech-to-speech models. You have opinions on the trade-offs, grounded in things you've actually built, not blog posts.
Strong software engineering. TypeScript/Node.js and/or Python, and the maturity to own a production service end-to-end: containers, cloud infrastructure, CI/CD, observability.
A latency obsession.