Screen-free curiosity
A voice-interactive AI agent that turns UK curriculum-aligned learning into imaginative, parent-led play, no screen required for the child.
Safety built in three layers deep
A bespoke safeguarding system, independently reviewed and praised for the level of protection it gives children interacting with generative AI.
Made the untestable testable
A virtual device simulator, cost model and AI evaluation engine let us develop, test and optimise the product before the hardware was ready.

Most children's screen time creeps in through the back door: a bedtime story on a tablet, homework help from a gamified app… it all becomes "just five more minutes" of something bright and loud.
Sandra, Founder of Jules Stories, set out to prove there's another way: a voice-activated device that children aged 4-10 could simply talk to, that could hold a story and teach them something, all without asking them to look at a screen.
The Jules Stories team had already built a basic hardware prototype, a WiFi-connected app and a proof of concept using Gemini's Live API. The core idea had legs. What they needed was a software partner to help turn it into a robust product. And that’s when Jules Stories joined us through our Impact Builder programme.
Jules Stories wasn't just about being anti-screen-time, it was about being pro-imagination. The idea was to build an AI agent, not a friend or a human substitute, that could hold a story-driven conversation with a child, stay inside a UK curriculum framework, and let parents set the learning objectives (such as reading, maths, and languages) their child would explore. Friendly, but not a friend, was the principle behind almost every decision on this project, from how the Jules device spoke to how it behaved.
That meant making an inherently unpredictable technology behave predictably, safely and affordably enough to put it in the hands of children. The AI agent needed:
Above everything, the child's safety was non-negotiable. Any product putting generative AI directly in front of young children has to answer some hard questions first: how do you stop an open-ended model from giving advice, discussing something harmful, or drifting somewhere a parent wouldn't want it to go? Every. Single. Time. As testing quickly showed us, a child’s curiosity to ask questions on-topic or not, doesn’t pause for a storyline.

We kicked things off with an in-person workshop with the Jules Stories team, as well as their own tech advisors, to do a technical deep-dive, reviewing the stack against the requirements the AI agent would need to meet. From there, the core architecture took shape: an AI storytelling agent (the "brain" behind the toy), built on Gemini and hosted on Google Cloud, alongside the backend and APIs connecting it to the parent app and physical device.
As the physical hardware was still being manufactured, we built a virtual testing tool to simulate the device in software, allowing development, safety testing and user testing to run in parallel with hardware production.
We added a RAG (Retrieval-Augmented Generation) layer that gives the agent the right UK curriculum context based on the child's age and school year, essentially acting as a reference library it can look up in real time rather than relying only on what the underlying model already knows. Voice mattered as much as words. We tested text-to-speech providers to find one that sounded warm and entertaining for young ears without tipping into sounding like a friend, then repeated the process for French and Spanish, balancing quality, cost and pronunciation across all three languages.

To make safeguarding as watertight as possible, we built multiple layers of protection around the agent, working before and after every exchange:
The result was a system designed to make an inherently unpredictable technology behave in a much more controlled and consistent way, without relying on any single layer to catch everything. But that thoroughness isn't free. Running these checks on every exchange increased latency, so we focused on getting response times as close to our target as possible without compromising security.
‘Jules is a model for what AI KidTech should be […] They underwent our highest tier of trajectory testing, with daily use across a variety of profile and behaviour archetypes for a significant length of time. We were checking against a range of potential harmful AI behaviour. Jules passed with flying colours.’
BrainSafe
Safety also had to hold up in three languages. Validating French and Spanish, languages the team wasn't fluent in, required us to lean even more heavily on the client, testing tools and safeguards we'd built rather than relying on our own interpretation of what safe sounded like.
AI projects don't behave like traditional builds, so we built the tools to match. Alongside the agent and virtual testing tool, we also developed:
Generative AI responses are probabilistic, so a single good response doesn't tell you much when the same prompt can behave differently next time. The evaluation tool automated a large chunk of testing to surface patterns at scale, but human review was still essential for deciding what good looked like.
It also meant treating behaviour tuning as an ongoing discipline rather than a box to tick early on. Small adjustments, made for good reason, could ripple through how the whole agent behaved.
Early on, the product started with an open-ended storytelling concept where Jules could tell stories freely. However, testing showed us that a more structured, turn-based experience would better support the learning objectives, as well as make the product’s running costs sustainable long-term.

The challenge was finding the balance between fun and learning: Jules needed to feel like the child's own adventure, while still functioning as a genuine learning tool, and it needed to give parents a say in what their child was learning without taking away the sense of autonomy that made it engaging.
Getting there meant asking some deceptively difficult questions: what actually makes a story good, and how much does that differ for a 4 year old versus a 10 year old? Our product and development team worked through it together, shaping pacing, structure and personalisation into the agent's storytelling logic, which was then further formalised through a storytelling framework.
Jules launched to MVP on time, with a product designed around the three things that had shaped the project from the start: safety, usability and cost.
Generative AI products don't stay still. The model, the costs and the behaviour all keep moving, so our approach had to be just as adaptive. Rather than treating each change as a wrong turn, we treated behaviour tuning and evaluation as a normal part of building something genuinely new.
Jules was never just one deliverable. Alongside the AI agent, our work touched the app experience, backend and APIs, as well as coordination with the manufacturing partner bringing the hardware to life. Post-launch, we worked with Jules Stories' new in-house developer to hand over the foundations and testing approach, helping set their team up for the next stage of development, including a wider rollout and broader user research through a dedicated agency.
The value of the project went beyond the MVP. We left the Jules Stories team with the tools, testing approach and foundations they needed to take the product into its next stage of growth.


Let's see what we can make together.
From startups to scaleups and enterprise, we're always happy to talk to impact-focussed and ambitious organisations. Use our quote creator to get a better idea of scope, timescales and next steps.