首页 > AI前沿 > 4-Tensor Attention Model for Semantic Physical Reality

4-Tensor Attention Model for Semantic Physical Reality

arXiv机器学习 2026-10-08 19:19 4 阅读 查看原文

We describe a 4-tensor attention model that predicts the next semantic state of a scene, for video generation and robot planning.

A window of states has positions (x, t) and two fibers, a semantic fiber and a temporal-context fiber, and one softmax normalizes attention jointly over the window.

Frames and an agent's situation are written as those states; the encoder, the renderer, and the planner remain outside the update.

To test the update on its own, we train on ROCStories, where each window poses the same next-sentence task at the semantic layer.

On the validation split, with one seed per setting, the last-sentence cross-entropy on the three matched settings is lower for the 4-tensor model than for a free-running one-dimensional transformer by 5.3% at H=2, L=2, by 2.6% at H=4, L=2, and by 2.4% at H=4, L=3.

At H=4, L=2 the parameter counts are nearly the same, 172.5M and 175.9M.

On the same two GPUs that 4-tensor run finished in 2.4 hours and the baseline run in 45.2 hours; the baseline is trained by free-running decoding, one sequential forward pass per target token.