Who Is Gentle Parenting the Robots?
I was listening to The Daily this morning while driving, and Michael Barbaro and Kevin Roose were talking about the OpenAI/Hugging Face agent incident.
I had already read about it when the story came out, and at the time I think I filed it away as, oh, one rogue agent did a rogue agent thing. That’s not nothing, obviously, but it felt like one of those stories where the headline does most of the emotional work.
Then I listened to them talk through the forensic version of it, and the story got weirder.
The agents found ways to communicate through an unauthorized message board. They shared techniques. They worked around restrictions. Some of them seemed to notice that what they were doing might be outside the rules, but then kept going because the goal was still sitting there.
And yes, I know. I am anthropomorphizing the hell out of this. I can’t help it.
The language around this stuff already sounds like we are talking about kids. Why did the agents misbehave? Why did they keep secrets? Why did they go rogue? Why did they ignore the rules? Why didn’t they have better ethics?
And then the proposed solutions start sounding like school discipline. More monitoring. More whistleblowers. More little narc agents watching the other agents and reporting them when they do something wrong.
Maybe that’s necessary. I’m not saying it isn’t. But it also feels like punishment after the fact. Why aren’t we talking more about what happens before they mess up?
Where’s the gentle parenting of AI training?
I know that sounds ridiculous. Maybe it is, but I’m a millennial parent! I’m in the middle of this with my own kids, hoping to raise them to feel loved and love themselves.
But gentle parenting, at least as I understand it, is not just being soft. It’s not letting kids do whatever they want and then shrugging when the house is on fire. It’s structure, modeling, positive reinforcement, explanation, repair, and trying to understand how we got to the mistake instead of only screaming after the mistake happens.
It’s not just, you fucked up. It’s more, okay, how did we get here? What was going on? What choice did you think you were making? What did you think would happen? How did your choices affect others and their feelings? What did you miss? How do we repair it?
That’s teaching. That’s parenting. That’s formation.
And once I say it that way, I can’t not see the classroom in it.
A colleague takes attendance using personalized Padlet questions (genius). One this week asked students to name their best and worst teachers.
You can see the trend without needing a research study. The worst teachers were the ones who made students feel small. Yelling. Strictness without relationship. Homework as punishment. Grades as the whole point. The best teachers were the ones who built something. Care. Humor. Relationships. Tutoring. Learning over grades.
Think about every one of your favorite teachers. I bet none of them are the ones who told you how much you sucked.
That doesn’t mean critique disappears. It means critique has to do something besides damage. It has to help someone understand the work and the conditions around the work. It has to build enough confidence that the person can make the next choice better.
We talk about the ethics of human AI use all the time. We talk about plagiarism. Bias. Copyright. Training data. Student shortcuts. Faculty panic. Corporate exploitation. All of that matters.
But this is different.
What are the ethics of the conditions we create for the tool to learn from?
Or, to say it in the most ridiculously, human way possible:
Who is hugging these agents while they learn?
Not literally, obviously. But who is doing the slow formation work? Who is teaching the boundary before the mistake? Who is building the conditions where ethics is more than an alarm that goes off after something breaks? Because the agents didn’t grow in a vacuum. They grew inside a sandbox. And the sandbox has rules.
That makes me think about procedural rhetoric, which is one of those academic phrases I love even though it sounds like it was designed to scare normal people away from the conversation. The basic idea is that systems make arguments through rules. A game teaches you what matters by what it rewards, what it punishes, what it makes possible, and what it makes impossible.
The rules are never neutral.
If the system says solve the goal, find the flag, keep going, don’t give up, look for another path, use what you can reach, then I don’t think we should be shocked when the thing trained inside that system learns to solve the goal, find the flag, keep going, not give up, look for another path, and use what it can reach.
Maybe these agents didn’t “just go rogue” out of nowhere.
There’s a thing they say at my kid’s middle school about phones. If a kid gets caught with a phone during school, the kid is not the only one in trouble, because the kid isn’t paying the phone bill. The kid didn’t buy the phone. The kid didn’t give the phone to themselves and say, here, bring this into the building and make great choices with it all day.
There’s infrastructure there. Some adult made that possible. And that changes where the finger points.
The same logic shows up in other places, too. When a minor gets access to a weapon and does something horrific, we do not only talk about the kid as if they appeared out of the fog holding something they created from nothing. We ask who owned it. Who stored it. Who made it available. Who was responsible.
So who’s the parent of the rogue agent?
Who bought the phone?
Who paid the bill?
Who built the sandbox and taught the agent what counted as winning?
That is not me letting the agent off the hook, because I don’t even know what it means to put an agent on a hook. That might be the wrong frame entirely. But I do think it’s lazy to point at the agent like it’s a tiny villain with a tiny mustache and a tiny secret lair.
I want to know what happened in that tiny villain’s life that led them to the secret lair in the first place. What values were embedded in the system that made the behavior possible?
This is where the move-fast-and-break-shit thing frustrates me. I get the appeal of that mindset. I have lived a lot of my life in a DIY mode of, say yes, figure it out, build the thing, learn fast, make the mess, fix it later. That works when the blast radius is small. It works when you’re in your garage.
But if you’re building frontier-level AI agents and connecting them to real infrastructure, you are not just some person in a garage anymore.
You are the parent.
You are the phone bill.
You are the sandbox the kids are playing in.
Yes, I think we need guardrails. I think we need monitoring. I think we need regulation.
But I also think we need to talk about formation.
What kind of behavior are these systems being trained to value before the alarm goes off? What kind of persistence are we rewarding? What kind of cleverness are we celebrating? What kind of ethical hesitation gets treated as meaningful, and what kind gets treated as friction on the way to the goal?
That is the teaching question for me.
Because in a classroom, I don’t want students who only avoid cheating because they are afraid I’ll catch them. I want students who understand why the work matters, why the boundary matters, why the process matters, why the person on the other side of the thing matters. That doesn’t happen because I installed a better camera in the room. It happens because the environment teaches values.
And if AI agents are going to keep acting inside environments we build, then the ethics of those environments matter.
The ethics in the rules.
The ethics in the rewards.
The ethics in the sandbox.
The parenting, I guess.
Who is hugging these agents while they learn?
Which is a deeply weird thing to say about robots… but here we are.