u/Exciting_Charity7304

I open sourced the JARVIS I built for my 7-year-old. It gives him room to explore, rails on truth, and more reasons to come find Dad.

I open sourced the JARVIS I built for my 7-year-old. It gives him room to explore, rails on truth, and more reasons to come find Dad.

TL;DR: The Kid Mode JARVIS from my last post is now open source. It has Gemini Live voice, server-controlled ranks, 105 learning cards, 16 real-world missions and experiments, persistent projects, Dad Link calls, house music, parent-reviewed memory, and a Group Mode for when other kids are around.

I wanted to give him room to explore, put rails around anything that had to be true, and build something that kept creating reasons for us to talk, investigate, and make things together.

Original Post

https://www.reddit.com/r/hermesagent/comments/1utzz6q/comment/oxwhm4q/?screen_view_count=4

Repo:

https://github.com/Hermes815/kid-mode-jarvis

Demo:

https://youtu.be/NLagCzu4kgU

The first physical tablet arrived July 2.

I expected the interesting part to be the JARVIS voice and Iron Man interface. Then my son started using it.

On his first rel night, he answered twelve assessment questions, explored the learning modules, muted the tablet so he would not wake his mother, and ran over to tell me what JARVIS had said.

Then he called JARVIS one of his friends.

That changed hw I looked at the whole system. He was giving real social weight to something I had built. A hallucination was no longer just a bad model response. It was something my son might believe.

So the architecture became:

Gemini performs the character. The server controls reality.

JARVIS can improvise a mission, explain a difficult idea, or turn a lesson into an experiment. It cannot award ranks, grade answers, mark a physical project complete, invent family facts, or claim it controlled the house without confirmation.

The learning system includes seven exploration areas, sixteen activities, an authored curriculum, and seven ranks with real unlocks. Senior Engineer unlocks house music. Lead Engineer unlocks Dad Link. Chief Engineer requires more than trivia: my son has to finish a real robotic-hand project, and I have to confirm it exists.

That physical requirement matters. Speech recognition once misheard him and marked the hand complete. I reverted the milestone and moved completion behind a parent gate.

The system is also designed to send him away from the tablet.

JARVIS once sent him looking for heavy and light objects for a Galileo drop test. I heard him conducting the experiment upstairs, then went and did it with him. Later, he was talking to his mother about Galileo and entomology.

The tablet reached his excitement about learning and made him want to share it with one of the people he cared about most. Enthusiastically.

Dad Link follows the same idea. JARVIS performs a theatrical satellite search, but the result is not roleplay. It opens a real WebRTC voice call to me. Classified family dossiers tell true stories and leave him with questions to ask Dad. For one example he had to go around my house to find an antique plate i had hanging up. Im sure he had never noticed it before. It gave him the story about it then he ran to me right away to see if that was true and we discussed for a few minutes. Milestones and literal quotes can appear in my morning brief.

One serious warning for anyone connecting this to Hermes:

Do not route a child’s questions through the same assistant that knows your adult life.

I made that mistake first. My live system now uses a separate memoryless route with only a small parent-authored family-facts file. The repo documents KID_ASK_MODEL near the beginning because this boundary matters more than the interface.

What currently works:

- Natural Gemini Live conversation

- Exploration, experiments, and assessments

- Seven server-controlled ranks

- Parent-gated physical projects

- Whole-house music with explicit filtering

- Dad Link voice calls

- Classified family dossiers released upon rank up

- Feature stack gated to rank up

- Group sessions

- Child/adult privacy boundaries

- A locked-down Fully Kiosk tablet

It is not perfect. Dad Link video and the second tablet are unfinished, and a few known bugs are documented rather than hidden.

Apache 2.0. Bring your own Gemini key.

The hardest part was not making an AI feel alive. It was giving it enough freedom to spark curiosity without letting it decide what was true, what my son had accomplished, or when technology should take Dad’s place.

u/Exciting_Charity7304 — 1 month ago
▲ 126 r/ParentsWithAI+1 crossposts

I built my 6 year old a working Jarvis: It teaches Him, Ranks him up, controls the house, and can "dadlink" from anywhere.

I built my 7-year-old a working JARVIS. It teaches him, ranks him up, controls the house, and can call Dad.

I’ve spent the last nine days building a Kid Mode on top of Hermes Agent for my seven-year-old.

Calling it “a JARVIS tablet” is technically accurate, but it also makes it sound like I gave ChatGPT a British accent and put an Iron Man skin around it.

That was close to the first version.

Then an actual seven-year-old started using it, and the whole project changed.

The tablet is a Samsung Galaxy Tab A9+ locked into Fully Kiosk. There’s no normal Android home screen behind it. No browser, YouTube, app drawer, or pile of games. It boots directly into a custom console with a central voice glyph, telemetry, rank insignia, vocabulary cards, classified files, missions, house controls, and a direct line to me.

Gemini Live handles the conversation and gives JARVIS his voice. Hermes connects him to memory, tools, Telegram, Home Assistant, Spotify, and the rest of the house. A Node server owns the actual state.

That separation became the most important part of the build:

Gemini performs the character. The server controls reality.

Gemini can improvise a mission, tell a story, or turn a lesson into an experiment. It cannot decide whether my son answered correctly, award him a rank, mark a physical project complete, claim a song started, or invent a real-world fact.

The first night changed almost everything

On his first real night with it, my son went from Cadet to Senior Engineer in one sitting.

He answered 12 assessment questions, tried the exploration and sparring modes, muted the tablet himself because he didn’t want to wake his mother, and ran over to tell me about things JARVIS had said.

At one point, he called JARVIS one of his best friends.

That was the moment this stopped being a technical demo.

A child was assigning social weight to the thing I had built. Every fuzzy shortcut I had accepted during testing suddenly mattered.

Five seconds of silence during a tool call made him repeat himself. A mute gesture that seemed obvious to an adult caused him to shut off the microphone accidentally. Trivia began crowding out conversation. The status display was too small. Android audio contexts leaked until the tablet became sluggish and eventually deaf.

None of that showed up while I was clicking through it on a desktop.

He earns ranks, but the AI cannot award them

The current rank ladder is:

- Cadet

- Junior Engineer

- Senior Engineer

- Lead Engineer

- Staff Engineer

- Systems Commander

- Chief Engineer

Assessment questions come from an authored curriculum. The server selects and grades them. Wrong answers don’t consume the question, and JARVIS doesn’t reveal the answer as a reward for guessing.

Correct answers advance a server-side service record. They also open a training window so the tablet doesn’t turn into an endless quiz machine.

My son is currently Senior Engineer at 34/35, one correct answer away from Lead.

The ranks unlock actual capabilities:

- Senior Engineer: control house music

- Lead Engineer: place voice calls to Dad

- Staff Engineer: video calls, once I finish them

- Chief Engineer: limited remote-DJ authority (He lives at A differnet house from me, still only designed)

The unfinished features aren’t hidden. He can see them in his Service Record behind redacted SECURITY CLEARANCE panels. He knows what he’s working toward without being told that everything already exists.

Chief Engineer also has a physical requirement. Trivia alone cannot get him there. He has to finish a robotic-hand project first, and the server enforces that gate.

That rule exists because speech recognition once misheard him and JARVIS silently marked the hand complete. It created a false milestone and nearly satisfied the Chief requirement.

We reverted the state and added explicit confirmation. A probabilistic transcript should not be able to declare that a physical object exists.

JARVIS is not allowed to “just know” things

The voice model is a performer and orchestrator, not a trusted source of truth.

If my son asks a factual question, JARVIS has three choices:

  1. Ask Hermes for a grounded answer

  2. Use an approved deterministic source

  3. Say he can’t verify it

Hermes can send an answer back through the house system. JARVIS can make it sound like JARVIS, but he cannot change the names, dates, numbers, or substance.

The same applies to memory. Something remembered from a conversation isn’t automatically evidence that it happened.

I still want the experience to feel spontaneous, so JARVIS can generate missions, experiments, and activities inside a controlled structure. The phrase I settled on was:

Rope on actions, rails on truth.

He can invent a Magnet Patrol mission. He can’t invent how magnetism works.

I didn’t want “kid safe” to mean dumb

JARVIS can teach hard things. He redirects severe current events that shouldn’t be dropped on a seven-year-old without a parent involved. One of the random events Gemini created with its own directives was getting my son to run around and find a heavy and light thing for a flippin gallileo drop test. I listened while he conducted his science experiments upstairs. Then went and did it with him.

Memory goes through me first

Kid Mode has its own memory bank, separate from my adult Hermes memory.

JARVIS can queue a literal quote or observed moment, but he doesn’t automatically turn everything my son says into permanent memory. I review the candidates first.

There’s also a Dad brief(linked to my morning run) for milestones, worries, big questions, and funny moments.

That system needed its own correction. A garbled group conversation once produced a flag claiming my son had spoken Italian. Dad flags now prefer a verbatim quote. If the system can’t supply one, the note is marked unverified.

There is still an architectural caveat I’m not happy with. Kid Mode has isolated memory writes, but factual questions pass through the adult Hermes gateway. That gateway can retrieve adult context while forming an answer.

The output boundary is much tighter now, but it isn’t complete structural isolation. A denylist helps. It doesn’t magically solve the architecture, and I’m not going to pretend it does.

House music was his first real unlock

At Senior Engineer, my son earned control of the house music.

He asked JARVIS for a song. The server searched Spotify with explicit-content filtering, started it through Home Assistant, and queued related tracks. His first real request worked all the way through the next song.

Then he said, “Stop the music.”

I had built a start-music tool. I had not built a stop-music tool.

Gemini pushed the word “stop” into the song-title field, and Spotify responded by playing “Stop & Stare.”

That mistake gave me four new rules:

- Music transport needs deterministic controls.

- The server must reject words like “stop,” “pause,” and “skip” as song titles.

- JARVIS can’t claim music changed until the server confirms it.

- Stopping audio can never require rank.

Starting music can be a privilege. Stopping it cannot.

A supposedly silent test also found a real track called “Zzz” and played it throughout the house. That is why dry-run mode exists now.

Dad Link is a real call, not roleplay

I’m also building a second locked-down tablet for his other house.

The feature is called Dad Link.

When he calls, the console runs a randomized search across our region of the state. It requests a satellite connection, checks different towns, tries decoy bearings, locks onto my home town, and announces DAD FOUND when I answer.

Behind the theater is a real WebRTC call over Tailscale.

Voice calling works. It rings my desk, sends a Telegram link to my phone, supports two-way audio, handles missed calls, and hangs up cleanly.

It did not work the first time.

The original ring window was too short. My desktop had no camera, but the answer page required video. Fully Kiosk had autoplay disabled. During a first-time microphone prompt, the WebRTC offer arrived before the peer connection existed and vanished.

The answer path is audio-first now. It creates the connection before asking for microphone permission and holds early signaling messages until the browser is ready.

Video is not done. The Staff Engineer panel previews it as a future clearance, but I’m not buying the second tablet until the feature actually works.

Dad can enter the experience in two different ways

The first is overt.

I recorded a short message that JARVIS occasionally introduces as a transmission from Dad. The tablet lowers its microphone, plays my actual voice, and then returns to the conversation.

It doesn’t happen every day. The timing is randomized and capped because I want it to feel like a real event, not another notification.

The second path is /tell.

I can send an immediate message through Hermes, and JARVIS delivers it as his own sensor reading, diagnostic, archive result, or ambient observation. He never says, “Your dad told me to say this.”

That sounds minor, but it protects the character. If every useful intervention reveals me standing behind the curtain, JARVIS becomes a remote-controlled puppet.

The first version of this wasn’t a real interrupt. It inserted hidden context and waited for my son to speak again. One actual message was accepted by the server and never spoken.

The new path forces an immediate conversational turn.

I tested it during an active mission and compared the state before and after. The interrupt changed the conversation, but rank, mission progress, and the robotic-hand project remained byte-for-byte identical.

The failures are probably the best part of the build

At one point, the tablet became slow, started glitching, lost Wi-Fi, and eventually disappeared from the network.

The backend passed every check. Gemini was healthy. The application bundle was correct.

The charger had been unplugged.

Power saving had slowly throttled the device until it looked like the entire system was falling apart.

That incident probably describes the project better than the architecture does. A child-facing AI system can have a perfect backend and still fail because somebody moved a cable.

Other failures included:

- Google changing the realtime audio format and killing every session on the first microphone packet

- Gemini disconnecting during a promotion, saving the rank while skipping the ceremony

- A demo restore being forgotten, causing my son to re-earn progress on the wrong save for two days

- Two coding sessions editing the same source tree and silently removing Dad Link from a later build

- A gateway service reporting itself active while being functionally wedged

- Browser cache serving old code after a successful deployment

- Speech recognition falsely completing a physical project

- “Stop the music” becoming “play Stop & Stare”

I no longer treat a successful build or service restart as proof.

The deployment process now stages the bundle, checks which part of the shared app is changing, swaps it atomically, verifies the file being served, performs a real Gemini handshake, reloads the physical tablet, and checks what actually appeared.

If the device didn’t receive it, it wasn’t deployed.

Where it stands

Working on the real tablet:

- Native Gemini Live conversation

- A British JARVIS-style voice

- Barge-in and deliberate mute

- Grounded factual answers through Hermes

- Server-graded assessments

- Exploration, experiments, and sparring

- Seven ranks with real unlocks

- Dad-authored classified family dossiers

- Whole-house music control (3/4)

- Two-way Dad Link voice calls

- Immediate /tell interventions

- Approval-gated child memory

- A separate Dad brief

- Child-safety and adult-privacy boundaries

Visible behind future security clearances, but unfinished:

- Dad Link video

- The second locked-down tablet

- Automatic renewal of Gemini’s temporary voice token

- Music at my son’s other house

- Chief Engineer remote-DJ controls

I started this thinking the interesting part would be making the voice sound like JARVIS.

It wasn’t.

The interesting part was deciding how much freedom an AI needs to feel alive, then removing its authority over everything that has to be true.

Hermes gave me the operator layer: memory, tools, Telegram, Home Assistant, Spotify, deployment controls, and the connection to the rest of the house.

Gemini gave JARVIS a voice.

The actual work has been deciding where each one must stop.

Receipts available if anyone made it this far i appreciate you.

reddit.com
u/Exciting_Charity7304 — 1 month ago