Building a GCSE study tutor, and taking it to production
My daughter was struggling with GCSE English, and a demo of a Reachy Mini robot running on a DGX Spark gave me the idea. It began as something for the house; the current work is deploying it properly.
My daughter was struggling with GCSE English, which is what started this off. Around the same time I watched a demo of Hugging Face’s Reachy Mini robot working with an NVIDIA DGX Spark, and the two came together into an idea: a study tutor she could talk to, running on a DGX Spark, with the robot on the desk as the thing she actually interacts with.
What it is
The tutor takes a Socratic approach, so it asks questions rather than dispensing answers, and if you ask it to simply tell you the answer it politely refuses. Behind it sits a fine-tuned Gemma 4 26B served through llama-swap, a retrieval store holding the three set texts, a student model in Postgres that remembers progress between sessions, and a gamification system built around kindness, with no leaderboards.
There are two ways to use it. A Flutter mobile app with tap-to-talk voice is how my daughter uses it day to day, and the Reachy Mini on the desk runs the same sessions alongside it, its spoken path currently being rebuilt onto the same streaming voice. Speech is streamed rather than batched, so a transcript comes back in around 0.2 seconds and the first spoken sentence in about three.
Two rules matter more than the rest: every quote is checked against the actual text before it is shown, spoken or saved, and exam board material is excluded entirely.
An honest evaluation result
After fine-tuning the model I ran a blind evaluation against the well-prompted base model, and the result wasn’t what I had hoped for. The base model met 88.5% of the expected tutoring behaviours; the fine-tune managed 62.5%, winning only on holding the Socratic stance. I published that as it stood in the Kaggle hackathon submission rather than re-running it until it looked better, and a fuller re-run in August told the same story.
Using it for real turned up other problems too, including model reasoning leaking into a live transcript and a bearer token that ended up in public repositories. Both are written up with dated fixes.
Taking it to production
Building it to run in the house suited the privacy question well, since a child’s learning data stays on a machine at home. The difficulty is that nearly every AI engineering job advert now asks for production deployment experience, and a system that only ever runs in my house doesn’t answer that. So the deployment became a deliverable in its own right.
The costed default is AWS: the same stack moved to a UK region on an EC2 g6.xlarge running llama-swap, at roughly $70–75 a month, starting with a one-evening spike costing about $5 to prove the shape. Bedrock’s Custom Model Import was a dead end for this model across three attempts. Moving a child’s data off a machine at home also changes the data protection question, so I settled that first: UK region, encryption at rest, consent captured during onboarding, an erasure path with a 30-day service level, and a DPIA.
AWS may not be the answer. Hugging Face’s own write-up points at NVIDIA Brev and Hugging Face Inference Endpoints, and Reachy Mini applications are pip-installable Hugging Face Spaces, which fits this project well. Nothing is deployed yet — the decisions are made and costed, and the next step is the spike.
Summary
The fine-tuned model has turned out to be the least important part of this; what the tutor can be trusted with comes mostly from the machinery around it. Building something that works on your own hardware is a good way to learn, but deploying it, governing the data and running it for other people are a separate discipline — and one the market is asking about far more loudly.