Senior AI Data Scientist — Agentic Process Development
14 ore fa
Magliano in Toscana, Tuscany, Italia
team.blue
Tempo pieno
Gratuito con email o Google
Salva questo lavoro e mantieni la tua ricerca organizzata
Crea un account gratuito per salvare lavori, creare avvisi e tornare a questa inserzione dalla tua dashboard.
Gratuito con email o Google
Company Overviewteam.blue is the market leader in enabling digital success for small and medium-sized businesses (SMBs) across Europe, catering to over 3 million customers in 25+ languages.
Scopra di più sui compiti quotidiani, le responsabilità generali e l'esperienza richiesta per questa opportunità scorrendo subito verso il basso.
Our mission is to make online business success simpler, by providing our customers with all the tools and resources they need to excel online and remain ahead of the curve.Position OverviewWe are looking for a Senior AI Data Scientist to streamline HR processes at team.blue — not by analysing them, but by building agentic systems to run them.
Recruitment, onboarding, performance, rewards and offboarding are each multi-step processes spanning several systems and up to 25 countries, and your mandate would be to create systems that can streamline them end to end.The method matters more than the domain: map a process, quantify what it costs in headcount, score which steps an agent could take, build a proof of concept, and take it to production.
This work sits closer to building autonomous, side-effecting systems than to building predictive models.
The agents you design would be able to revoke IT access, issue signed contracts, and flag pay outliers into approval workflows.
A wrong output here is not a bad number someone can catch — it is a high impact action taken in the world.What we are actually screening for:Not whether you can hand-roll a gradient-boosted tree.
LLM coding tools can do that faster than you can.
Classical ML and applied statistics are the entry fee for this role — necessary, and assumed.
Everyone we are talking to has them.What separates candidates is whether you can build an agent that is robust, cost-effective and trustworthy — with deterministic operations rather than “LLM does everything” patterns.
Building a demo is now easy.
Knowing whether to trust one is not.We also mean end to end literally.
You write it, you containerise it, you instrument it, and you own it when it breaks.Your day would involve:Time with the HR Ops lead mapping how a leaver actually gets offboarded across 19 countries — then turning that into a process inventory with FTE cost attached per stepFacilitating a half-day session with Talent, Rewards and HR Ops leads to score automation candidates on impact, feasibility and LLM/tool fit — extracting requirements live from people who do not think in data modelsDesigning the state transitions: what triggers, what branches, which systems get called, where it waits, when it escalates, and what happens when step 4 of 9 failsBuilding the guardrails before the capability — dry-run mode, an approval gate ahead of anything irreversible, least-privilege scoped credentials, a rollback pathDeciding where a human stays in the loop, at what confidence threshold, and designing a review queue they will actually useWriting evals for output that precision and recall do not capture: task-completion rate, hallucination rate, gendered or culturally biased language in AI-drafted reviewsWiring an agent to a webhook instead of a nightly batch pull — and making the handler idempotent so a retry does not offboard someone twiceDeciding which steps in a flow warrant a frontier model and which can run on something cheap, then proving that routing decision with numbersSitting in a vendor demo asking what their API actually exposes, what their data model looks like, and what integration really costs usWhat you will bring:7+ years building data and ML systems in industry, spanning both sides of the LLM shift.
We want the judgment that comes from having debugged systems before you could ask a model what was wrong.Somewhere in that history: you have shipped something that had permission to take an irreversible action affecting real customers — and you can tell us what you did to sleep at night.Expert in Python and ML.You ship end to end.
Python someone else can still read in six months, a current toolchain (uv, Docker or an equivalent — we care that you re-examine your tooling, not which tool you landed on), your own container, your own instrumentation.Production experience with multi-step, tool-calling LLM workflows — orchestration, retries, idempotency, timeouts, partial-failure recovery.
State-machine design, not only train/serve pipelines.Cost and latency engineering as a first-class concern — model routing, caching, batching, and the instinct to know what a flow costs per run before Finance asks.A safety instinct for systems that take actions — staging modes, approval gates, least-privilege scoping, rollback.Evaluation design for generative and agentic output — LLM-as-judge, golden-transcript regression suites, red-teaming.Applied statistics you can adjudicate with.
Not "can build a model" but can tell us whether a number is trustworthy and what would have to
Scopra di più sui compiti quotidiani, le responsabilità generali e l'esperienza richiesta per questa opportunità scorrendo subito verso il basso.
Our mission is to make online business success simpler, by providing our customers with all the tools and resources they need to excel online and remain ahead of the curve.Position OverviewWe are looking for a Senior AI Data Scientist to streamline HR processes at team.blue — not by analysing them, but by building agentic systems to run them.
Recruitment, onboarding, performance, rewards and offboarding are each multi-step processes spanning several systems and up to 25 countries, and your mandate would be to create systems that can streamline them end to end.The method matters more than the domain: map a process, quantify what it costs in headcount, score which steps an agent could take, build a proof of concept, and take it to production.
This work sits closer to building autonomous, side-effecting systems than to building predictive models.
The agents you design would be able to revoke IT access, issue signed contracts, and flag pay outliers into approval workflows.
A wrong output here is not a bad number someone can catch — it is a high impact action taken in the world.What we are actually screening for:Not whether you can hand-roll a gradient-boosted tree.
LLM coding tools can do that faster than you can.
Classical ML and applied statistics are the entry fee for this role — necessary, and assumed.
Everyone we are talking to has them.What separates candidates is whether you can build an agent that is robust, cost-effective and trustworthy — with deterministic operations rather than “LLM does everything” patterns.
Building a demo is now easy.
Knowing whether to trust one is not.We also mean end to end literally.
You write it, you containerise it, you instrument it, and you own it when it breaks.Your day would involve:Time with the HR Ops lead mapping how a leaver actually gets offboarded across 19 countries — then turning that into a process inventory with FTE cost attached per stepFacilitating a half-day session with Talent, Rewards and HR Ops leads to score automation candidates on impact, feasibility and LLM/tool fit — extracting requirements live from people who do not think in data modelsDesigning the state transitions: what triggers, what branches, which systems get called, where it waits, when it escalates, and what happens when step 4 of 9 failsBuilding the guardrails before the capability — dry-run mode, an approval gate ahead of anything irreversible, least-privilege scoped credentials, a rollback pathDeciding where a human stays in the loop, at what confidence threshold, and designing a review queue they will actually useWriting evals for output that precision and recall do not capture: task-completion rate, hallucination rate, gendered or culturally biased language in AI-drafted reviewsWiring an agent to a webhook instead of a nightly batch pull — and making the handler idempotent so a retry does not offboard someone twiceDeciding which steps in a flow warrant a frontier model and which can run on something cheap, then proving that routing decision with numbersSitting in a vendor demo asking what their API actually exposes, what their data model looks like, and what integration really costs usWhat you will bring:7+ years building data and ML systems in industry, spanning both sides of the LLM shift.
We want the judgment that comes from having debugged systems before you could ask a model what was wrong.Somewhere in that history: you have shipped something that had permission to take an irreversible action affecting real customers — and you can tell us what you did to sleep at night.Expert in Python and ML.You ship end to end.
Python someone else can still read in six months, a current toolchain (uv, Docker or an equivalent — we care that you re-examine your tooling, not which tool you landed on), your own container, your own instrumentation.Production experience with multi-step, tool-calling LLM workflows — orchestration, retries, idempotency, timeouts, partial-failure recovery.
State-machine design, not only train/serve pipelines.Cost and latency engineering as a first-class concern — model routing, caching, batching, and the instinct to know what a flow costs per run before Finance asks.A safety instinct for systems that take actions — staging modes, approval gates, least-privilege scoping, rollback.Evaluation design for generative and agentic output — LLM-as-judge, golden-transcript regression suites, red-teaming.Applied statistics you can adjudicate with.
Not "can build a model" but can tell us whether a number is trustworthy and what would have to