An intuition unconsciously acquired through living in the natural world. A pre-verbal spatial awareness of which structures fit together and where a mismatch causes things to jam. A while back, I came across an example of how this kind of awareness, which is second nature to living beings including humans constantly facing the real world, can be notoriously difficult for machines that learn from text and data to grasp.

It happened when I was working on a task to generate a 3D model of an early cryptographic machine along with the openable wooden box that houses it. Judging by how the main body of the box and the lid were connected, our basic spatial sense tells us that, no matter how you look at it, the structure makes it impossible to close. Humans, or certain other living beings, understand this instantly without even thinking. But because AI has never lived in the real world, it may have simply concluded that as long as the numbers and procedures lined up, the job was done. The model was GPT-5.5 from May 2026.

Now, here is where the real challenge began. How do you convey this absurdity to a GPT that lacks lived experience in the physical world? A casual nudge like, "If you connect it there, it won't close," didn't improve things at all. Next, I suggested, "Attach a hinge to the edge of the hollowed-out side of the lid, and connect it to the edge on the main body side," but that only resulted in a hinge being awkwardly tacked onto the lid, doing nothing to fix the fact that it still couldn't close. When you actually try to translate everyday common sense that we usually take completely for granted into an expression a machine can understand, you find yourself wondering where on earth to break through.

The explanation I finally came up with was this:

The lid is still connected on the side that cannot be closed. I want to make sure, as I think a fundamental misunderstanding is happening: when you combine two hollowed-out rectangular cuboids to make a box that can be closed, you join the edges of the hollowed-out faces with a hinge, right? The current lid is being joined on the side that isn't hollowed out.

After waiting a few minutes, a closable box infused with a sense of the natural world was finally completed. It was a fascinating experience that reminded me how the common sense we humans share doesn't always align with a machine's processing patterns.

Of course, labs like Google DeepMind are putting a lot of effort into physical space simulation, and as learning data accumulates around Physical AI such as humanoid robotics, these kinds of inconsistencies will be eradicated at an accelerated pace, just like precedent cases with language models. Even so, I am certain there will still be times when what is obvious to a living organism remains completely unknown or unexperienced to an AI, accidentally leaving it outside the scope of consideration.

The trend toward leaving wide-ranging tasks and decision-making up to AI is probably unstoppable by now, but it really made me deeply feel the necessity of occasionally returning to lived experience, and exercising caution in areas where the damage from misalignments could be severe.