SEPTEMBER 2026 · RESEARCH CONCEPT · v4

Physical Firmware

A structured world model for reusable physical intelligence.

A robot should not need to relearn the reusable structure of physical interaction whenever its task, objects, or body change. Physical Firmware asks whether a learned model of action consequences transfers better when selected geometry, mechanics, contact constraints, and uncertainty are explicit parts of its architecture.

Status: research hypothesis / proposed architecture.
Not an experimentally validated system. No Physical Firmware performance results are claimed here.

01 / A physical prediction layer

Sensors → state / beliefVision, joint state, force, touch, history
Embodiment adapterGround state; model how commands become motion
PHYSICAL FIRMWAREBelief + candidate actions → possible futures
Predictive distributionMotion, forces, contact modes, uncertainty
Planner / VLA / task policyEvaluate futures; choose a command
Controller → actuation → world
Figure 1. Proposed interfaces, not a compulsory serial stack. The adapter participates in state estimation; the planner repeatedly queries the model before acting. Hardware protection remains independent.
Read the main text for intuition.Open technical notes for equations and assumptions.Read the case against it.

01 / THE PROBLEMKnowing what to do is not the same as predicting what will happen.

A robot pushes a part toward a fixture. Change the surface finish, the hidden weight, or the controller delay. The instruction stays the same. The consequences do not.

A policy maps observations and goals to actions. A world model predicts consequences, usually conditioned on candidate actions. A vision-language-action model, or VLA, is a policy that connects visual observations and language to robot actions. Modern VLAs already contain substantial implicit knowledge of physical interaction; generalist behavior and cross-robot transfer are demonstrated research directions, not capabilities this proposal invents.[1][7]

What remains difficult is learning reliably from interventions in an unforgiving world. Physical trials consume robot time, resets, operator attention, and sometimes hardware. Rare contact failures matter even when average performance is good. A camera may not reveal mass, center of mass, friction, or compliance. Small geometry errors can change a collision into a near miss; small trajectory errors can compound until a plan enters an unfamiliar state.

Contact also changes the kind of motion that is possible. Before touching a surface, a gripper moves freely. After touching, the same command may produce force, deformation, sticking, or slip. A policy changes the distribution of states it encounters as it improves—or as it makes a mistake. Prediction evaluated only on recorded demonstrations can therefore misrepresent prediction under the policy's own actions.

Learning behavior means learning useful choices. Learning dynamics means learning how interventions change the world. The two can share representations and training objectives. The unresolved question is whether selected explicit physical structure is a better inductive bias—a restriction or preference built into learning—for transfer, prediction, and control.

HYPOTHESIS

Some knowledge of physical interaction may survive changes of task, object, and embodiment more reliably when its predictive model is structured around entities, geometry, symmetries, constraints, hybrid events, and uncertainty.

02 / THE HYPOTHESISTrain broadly. Adapt lightly. Learn specifically.

Physical Firmware is a reusable, action-conditioned, probabilistic world model for physical interaction, with architectural structure derived from mechanics and learned residual dynamics derived from data. “Action-conditioned” means predictions depend on what the robot might do. “Probabilistic” means the output represents multiple plausible outcomes, not just one guessed trajectory. A residual is a learned correction for effects the chosen mechanics does not explain.

The computer analogy—hardware → firmware → operating system → application—is about reuse, not literal software placement. A physical prediction layer would be queried by task intelligence and grounded by robot-specific interfaces. It need not run on a microcontroller, live below the operating system, or own the control loop.

02 / What might transfer?

Train broadly

Shared physical core

Interaction operators; geometric structure; appropriate conservation and dissipation; contact mechanisms; predictive uncertainty.

Adapt lightly — target

Embodiment grounding

Kinematics, morphology, sensors, actuator dynamics, compliance, gripper geometry, latency, and robot-specific residuals.

Learn specifically

Task intelligence

Instruction meaning, goals, costs, sequencing, behavior selection, VLA reasoning, and task policies.

Shared core + grounded embodiment + task objective → a usable robot system.

Figure 2. Three proposed scopes of reuse. Their boundaries are hypotheses: task relevance changes perception, and embodiment changes which physical interactions are reachable.

The intended benefit is to learn “how to achieve this outcome” while reusing knowledge of consequences. But “train once” would be misleading. New material classes, sensor configurations, or mechanisms may require core updates. Even within rigid manipulation, an adapter might become so large that the factorization stops being useful.

What is architectural, and what is learned?

Candidate structureLearned or inferred quantitiesWhat it does not guarantee
Object graph and shared interactionsEntity features, geometry, interaction parameters, evolving edgesCorrect object segmentation or composition outside training
Equivariant geometric operationsScalar functions and compatible vector/tensor featuresSymmetry of a scene with fixed gravity and fixtures
Energy, interconnection, dissipationEnergy function, inertia, damping, bounded residualsTrue energy, stability with arbitrary residuals, or accurate contact
Contact constraints / hybrid modesContact geometry, friction, mode probabilities, compliant correctionsA correct rigid-contact approximation for every material
Probabilistic rollout interfaceState and parameter posteriors, predictive distributionsCalibrated uncertainty or safe decisions

Proposed design: begin with a small selection of these components, then test each. A single monolithic network containing every physical formalism would obscure which assumption helps. The possible contribution is a useful combination and an experimentally supported reuse boundary; none of the individual ingredients is new.

03 / STATE UNDER PARTIAL OBSERVABILITYThe robot sees evidence, not the state of the world.

A Markov state contains enough information that the future depends on the present state and action, rather than the entire past. A single image is generally not one. Motion can be hidden, an object can be occluded, and friction can remain unknown until the robot acts. A belief is a distribution over possible hidden states given the interaction history.

The proposed state includes robot configuration, object poses and velocities, geometry, persistent identity, articulated structure, support relations, and a contact graph. It may also need mass and inertia estimates, friction, compliance, deformability descriptors, actuator state, and uncertainty. Some variables are slow parameters; others change quickly. Heat, wear, damage, or a shifting load can invalidate an assumption that parameters stay fixed.

The model does not need to reconstruct everything visible. It needs a decision-sufficient physical state: enough information to rank actions for the relevant task family. A box's printed logo may be irrelevant to pushing. The offset of its center of mass may be decisive. The difficulty is that the next task can make a previously irrelevant detail important.

Technical note: a belief, a transition, and an observation model

Let \(o_t\) be a synchronized multimodal observation at discrete time \(t\), \(a_t\) the commanded action over the next interval, \(z_t\) a latent physical state, \(\theta_t\) hidden physical parameters, and \(c_t\) a discrete contact or hybrid mode. Then:

\[b_t = p(z_t,\theta_t,c_t\mid o_{\leq t},a_{<t}).\]

A practical filter approximates this posterior with particles, a mixture distribution, or a recurrent latent representation with probabilistic heads. Given a previous belief, it first propagates through an action-conditioned transition, then updates with a likelihood for the new observation. Denote the combined hidden state by \(s_t=(z_t,\theta_t,c_t)\):

\[b_{t+1}(s')\propto p_\psi(o_{t+1}\mid s')\int p_\phi(s'\mid s,a_t,\theta_e)b_t(s)\,ds.\]

Here \(p_\phi\) is the shared transition with learned parameters \(\phi\), \(p_\psi\) is an observation model with parameters \(\psi\), and \(\theta_e\) describes the embodiment. The integral includes summation over discrete modes; the proportionality constant normalizes the belief. This equation assumes the augmented state is Markov and observations are conditionally independent of the past given that state. Unknown actuator delay must therefore appear in the augmented state or action history.

A learned encoder is not automatically a Bayesian filter, and a latent variance is not automatically physical uncertainty. Without an identifiable observation model, different latent states may explain the same observations. Training must connect the proposed belief to measured future geometry, forces, events, or other decision-relevant targets.

Representation is a tradeoff, not a ladder of sophistication

RepresentationWhat it makes easierWhat it can hide
Pixels / videoLearning from broad visual data; predicting appearanceForces, hidden properties, causal action alignment
Compact learned latentFast rollout and task-relevant compressionUnits, constraints, identity, state sufficiency
Point cloud / 3D / 4D geometryShape, spatial relations, geometry through timeOccluded surfaces, material properties, contact forces
Objects / scene graphIdentity, relations, variable object countSegmentation errors; distributed deformation
Explicit physical variablesUnits, mechanics, interpretable constraintsUnmodeled effects; expensive or ambiguous identification

Latent planning is established, while recent surveys organize manipulation models across visual, latent, geometric, and physical representations.[19][14] A plausible starting point is mixed: explicit poses and velocities for rigid objects, learned shape and material descriptors, and a recurrent belief for unobserved variables. The research question is how small that state can be before its compression destroys useful counterfactual predictions.

04 / OBJECTS & RELATIONSReuse the interaction, not the whole scene.

A dynamic interaction graph represents entities as nodes and possible physical relations as edges. A node can be a robot link, rigid object, fixture, or a set of particles representing a deformable body. Edges encode relative pose, proximity, contact mode, attachments, or constraints. An edge can appear when objects approach and disappear when they separate.

Applying the same learned interaction operator to object–table, gripper–object, and object–fixture pairs creates an opportunity for compositional transfer. Graph-based learned simulators demonstrate the usefulness of shared local computations for physical systems.[27] They do not establish that a model trained on a few contacts will generalize to arbitrary clutter or new embodiments.

03 / One scene, several kinds of relation

Scene interaction graphA robot link attaches to a gripper. The gripper contacts a part supported by a table. A fixture is fixed to the table. A dashed edge represents possible part-fixture contact. Robot linkGripperPartFixtureTable attachmentcontactsupportfixedpossible

Node

Pose · velocity · geometry · material descriptors · belief about hidden properties.

Edge

Relative geometry · potential contact · active mode · constraint · interaction features.

A new part reuses operators. Its geometry and parameters still need grounding.

Figure 3. Solid and dashed edges distinguish current relations from candidate contact. Local messages must still transmit the effects of a coupled scene.
Technical note: local messages do not make contact local
\[m_i=\sum_{j\in\mathcal N(i)}M_\phi(z_i,z_j,e_{ij}),\qquad \dot z_i=D_\phi(z_i,m_i,u_i).\]

Here \(z_i\) is node \(i\)'s state, \(e_{ij}\) its edge features with neighbor \(j\), \(\mathcal N(i)\) the neighbor set, \(M_\phi\) a shared message function, \(m_i\) aggregated interaction features, \(D_\phi\) a shared dynamics function, and \(u_i\) the input acting on the node. This is a candidate smooth-mode computation, not a complete contact solver. Summation is invariant to neighbor ordering; that alone enforces neither force balance nor momentum conservation.

Simultaneous contacts can couple distant nodes through a rigid chain. A shallow message-passing network may not resolve that coupling. More iterations, multiscale communication, or an explicit constraint solve may be necessary. Contact graph mistakes can dominate all improvements in the dynamics operator. Compare oracle graphs against inferred graphs to expose this failure mode.

05 / STRUCTURED CONTINUOUS DYNAMICSMechanics supplies useful constraints. It does not supply the whole model.

A neural ordinary differential equation learns a continuous-time rule for how state changes.[22] One can restrict that rule using mechanics. Hamiltonian networks derive conservative motion from an energy-like scalar; Lagrangian networks provide an alternative based on positions and velocities.[23][24] Neither is a sufficient description of all manipulation.

Proposed design: use a port-Hamiltonian-style substrate where smooth mechanical state is meaningful. “Port” means a channel through which energy enters or leaves, such as an actuator. Represent energy exchange, damping, and inputs explicitly; reserve residual modules for inadequately modeled effects. Use a separate treatment for nonsmooth contact and a different state or formalism where deformation demands it.

05 / A candidate smooth-mode dynamics block

Energy H

Learned or partly analytic kinetic and potential structure.

Interconnection J

Skew-symmetric coupling; no energy creation by this term.

Dissipation R

Positive-semidefinite loss, within the chosen model.

Actuation Gu

Realized input from the embodiment model.

Residual fres

Learned discrepancy; may invalidate energy guarantees.

Integrator + contact

Advance smooth state; resolve constraints and mode changes.

State and inferred parameters → structured vector field → numerical update → next state.

Figure 5. These are candidate modeling components, not six mandatory neural networks. Impacts are not hidden inside a smooth force term.
Technical note: Hamiltonian structure, passivity, and the residual problem

For a smooth conservative mechanical system in canonical coordinates, let \(q\) be generalized position, \(p\) conjugate momentum, \(x=(q,p)\), \(T\) kinetic energy, and \(V\) potential energy:

\[H(q,p)=T(q,p)+V(q),\qquad \dot q=\frac{\partial H}{\partial p},\quad \dot p=-\frac{\partial H}{\partial q}.\]

This assumes a valid canonical state and an autonomous, differentiable Hamiltonian. Constant energy does not imply a closed orbit, a correct trajectory, or stability. A learned \(H\) need not be the true physical energy, especially when its coordinates are latent.

\[\dot x=[J(x)-R(x)]\nabla H(x)+G(x)u+f_{\mathrm{res}}(x,u,z).\]

Here \(J=-J^\top\) is skew-symmetric interconnection; \(R=R^\top\succeq0\) is dissipation; \(\nabla H\) is the state gradient; \(G\) maps realized actuator input \(u\) into state dynamics; and \(f_{\mathrm{res}}\) is a residual that may depend on additional latent context \(z\). These quantities can also depend on inferred physical parameters and mode, suppressed for readability. For a time-independent \(H\), the chain rule gives:

\[\dot H=-\nabla H^\top R\nabla H+y^\top u+\nabla H^\top f_{\mathrm{res}},\qquad y=G^\top\nabla H.\]

The output \(y\) is conjugate to \(u\), so \(y^\top u\) is input power in a physically grounded realization. Skew-symmetry removes the \(J\) contribution. With no residual, the smooth dynamics cannot increase \(H\) except through the input port. That is an energy-balance property, not an unconditional stability theorem. With an arbitrary residual, even that statement no longer holds.

One option is to constrain residual power or place uncertainty around it; another is to accept loss of passivity and measure the consequences. Positive-semidefinite parameterization of \(R\), for example through a matrix factor, is architectural enforcement. A penalty encouraging it is only soft regularization. Nonnegative storage, appropriate equilibrium structure, and further conditions are needed for stability claims. Existing stable port-Hamiltonian work proves guarantees for its specified architectures; those guarantees do not automatically extend to this hybrid, partially observed proposal.[26]

A skew matrix alone also does not make an arbitrary interconnection a Poisson structure: additional identities are needed if that geometric claim is made. Rotations live on a manifold, not unconstrained Euclidean coordinates. Lie-group models preserve the appropriate configuration structure; Lagrangian descriptions may be more natural when velocities are observed and canonical momenta are not.[25] Explicit time dependence adds another term to the energy balance, and impacts require a separate jump balance. Updating inferred parameters can also change the model's energy estimate; that bookkeeping change must not be mistaken for physical work.

This is where an attractive physical prior can become a liability. A rigid state cannot explain a buckling package. A passive model cannot explain an omitted powered mechanism. If residuals learn almost everything, the energy scaffold may add cost without useful restriction. The relevant comparison is with a capacity- and data-matched model, not with an intentionally weak predictor.

06 / HYBRID CONTACTThe difficult physics happens when the rules change.

A hybrid dynamical system combines continuous evolution with discrete changes of mode. A part moves freely, hits a fixture, slides along it, sticks, then loses support. A grasp closing is an actuator event; a stable grasp additionally requires appropriate contact and force. An insertion can move from unconstrained motion to several simultaneous constraints.

04 / Contact is a branching process

Free motion
Impact
Persistent contact
↳ Stick ⇄ slipSwitch when forces and relative motion change
↳ Release → free motionSupport or normal force disappears

Guard: when does a transition occur?
Reset: what changes across that event?
Mode dynamics: what happens between events?

Figure 4. A simplified mode diagram. Real multi-contact scenes combine many such relations; there is no single universal sequence.

The challenge is not merely predicting that two shapes touch. Contact timing depends on small geometric differences. Sticking and slipping can be visually ambiguous. Rare impacts receive little training coverage. Multiple contacts transmit forces through an assembly, while stiff or deformable interactions couple fast local changes to slower motion. A predicted average between “slips” and “sticks” may correspond to neither physically possible outcome.

Technical note: guards, resets, complementarity, and friction
\[\dot x=f_c(x,u),\qquad g_{c\to c'}(x)=0,\qquad x^+=\Delta_{c\to c'}(x^-).\]

Mode \(c\) selects the within-mode dynamics \(f_c\). A guard function \(g_{c\to c'}\), together with its crossing direction and feasibility conditions, signals transition to mode \(c'\). The reset map \(\Delta\) takes the state immediately before the event, \(x^-\), to immediately after, \(x^+\). At a rigid impact, positions are usually continuous while momentum jumps. A guard can also depend on input, time, or hidden parameters; the compact notation omits these possibilities.

\[\varphi(q)\geq0,\qquad \lambda_n\geq0,\qquad \varphi(q)\lambda_n=0.\]

For an ideal rigid, nonadhesive unilateral contact, \(\varphi(q)\) is signed gap and \(\lambda_n\) the compressive normal contact force. Bodies cannot overlap; the surface cannot pull them toward itself; positive force requires zero gap. Zero gap does not require positive force. These position-level conditions are necessary, not a complete dynamics law: velocities, accelerations, other constraints, and an impact law determine the force or impulse.

\[\|\lambda_t\|\leq\mu\lambda_n.\]

Here \(\lambda_t\) is tangential contact force and \(\mu\geq0\) the Coulomb friction coefficient. The cone bounds admissible force. Sliding additionally needs a direction opposing relative tangential motion, while sticking requires compatible zero slip. At impacts, impulse variables and a restitution or dissipation rule replace an ordinary finite-force update. Coulomb friction neglects many effects, including adhesion, hysteresis, and velocity-dependent materials.

Rigid-contact complementarity can be nonunique or ill-conditioned. Time-stepping schemes provide an alternative to resolving every impact as a separate event.[29] Smooth compliant contact eases some optimization problems but introduces stiffness and approximation error. Differentiating through contact may require implicit differentiation, event sensitivity, or a surrogate; gradients can be unreliable exactly at mode boundaries. Choose the approximation for the task and report constraint violations.

Recent work changes the baseline

ContactWorld studies representations for vision-tactile predictive planning; its revised 2026 preprint emphasizes spatial structure, temporal continuity, and compatible modalities. Its lesson is that adding a tactile channel alone is insufficient.[16] Dream-Tac jointly models actions, future visual observations, and tactile dynamics, using contact-aware fusion.[18] Both belong in the comparison set for contact prediction, rather than being treated as peripheral perception work.

WorldContact, a September 2026 preprint, addresses deformable-object interaction and uses its learned dynamics to expand policy-training data.[17] This demonstrates a relevant architectural route: a predictive model can help a VLA during training without becoming its runtime planner. These papers support investigating contact-aware representations; they do not validate the complete Physical Firmware factorization.

07 / THE EMBODIMENT ADAPTERA small identification problem—if the hypothesis holds.

System identification infers a system's dynamics or parameters from input–output observations. The adapter should perform identification as well as encoding. It connects coordinate frames, sensor calibration, morphology, action interfaces, and the physical model. A position target, velocity command, and torque command are not interchangeable interventions.

The shared model has learned parameters φ. The robot has embodiment variables θe: kinematic structure, actuator gains and lag, joint friction, compliance, gripper geometry, sensor transforms, control frequency, and small residual modules. A morphology description can supply known structure; controlled interactions infer uncertain quantities. The desired result is to update θe while initially holding φ fixed.

Amortized identification trains an inference model across many systems so a new interaction history can produce a parameter estimate without fitting everything from scratch. Rapid Motor Adaptation is a relevant example of history-based online adaptation in locomotion; it is not evidence that a manipulation predictor will transfer across arbitrary robots.[31]

Calibration should ask identifiable questions

Free-space motion at different speeds can reveal lag and actuator response. Gentle surface contact can ground geometry and compliance. Pushing, lifting known loads, opening and closing the gripper, controlled slip, and bounded force application reveal different combinations of parameters. Calibration must obey the robot's validated limits and use force-limited procedures appropriate to its hardware.

Not every parameter is separately identifiable. A slow push may reveal a combined friction effect but say little about inertia; actuator tracking error can masquerade as an object-dynamics error. Keep parameter uncertainty when multiple explanations remain consistent. Choose additional probing actions for information only when their expected benefit justifies their risk and time.

Technical note: commands are not physical inputs
\[h_{t+1}=A_{\theta_e}(h_t,a_t),\qquad u_t=B_{\theta_e}(h_t,a_t).\]

Here \(a_t\) is a command, \(h_t\) is actuator/controller history including delayed commands, \(A_{\theta_e}\) updates that history, and \(B_{\theta_e}\) predicts the physical input \(u_t\) realized over the next interval. The dynamics core consumes that realized input or its distribution. A memoryless map is appropriate only when lag and controller state can be neglected at the modeling time scale.

A posterior over \(\theta_e\) can be inferred from calibration observations and carried into rollout. A robot-specific residual should be capacity-limited and audited. If the adapter learns the entire task or absorbs the whole world model, calling the remaining core “shared” does not demonstrate useful transfer. Report adapter capacity, data, compute, and retained performance on previous embodiments.

RESEARCH QUESTION

Can a new embodiment be characterized with a small amount of structured interaction, at lower total cost than adapting a strong generalist policy or retraining the predictor?

08 / UNCERTAINTYA distribution over futures, not a confidence badge.

A robot can be uncertain because it cannot see the contact, because the friction is unknown, or because its model has never encountered this material. Those are different reasons to hesitate. They may call for a better view, a probing interaction, a shorter planning horizon, a conservative action, or a handoff.

UncertaintyMeaningUseful response
AleatoricStochasticity or ambiguity remaining under the chosen sensing and state representationPropagate possible outcomes; improve sensing when hidden state is recoverable
EpistemicLack of knowledge in the learned modelDetect unfamiliar conditions; acquire informative data
ModeWhether an impact, stick, slip, or release will occurKeep distinct branches instead of averaging them
ParameterUnknown mass, friction, compliance, or actuator responseMaintain and update a parameter belief
HorizonHow predictive reliability changes with look-ahead timeEvaluate calibration separately at each useful horizon

These are overlapping axes, not mutually exclusive error bins. Hidden friction creates parameter uncertainty and may change the contact mode. “Irreducible” aleatoric uncertainty is relative to the available information: adding tactile sensing can make previously hidden variation predictable.

06 / One action, several plausible consequences

Branching probabilistic futuresFrom a current belief, an action produces branches that settle, slide, or lose support. The spread increases with horizon. This is a qualitative schematic, with no probabilities or empirical measurements. NowMode ambiguitySlideSettleLose support Prediction horizon → · conceptual, not measured Possible futures, mobile diagramAn initially narrow set of predictions branches into sliding, settling, and loss of support. This is conceptual, not a measured distribution. NowSlideSettleLose supportLonger prediction horizon →
Figure 6. Branches illustrate alternative contact outcomes, not a calibrated prediction region. A mean trajectory can conceal a failure branch. Spread often grows with horizon, but contraction or a predictable constraint can reduce it.

Probabilistic ensembles with trajectory sampling, as in PETS, offer a practical starting point.[20] An ensemble represents disagreement across fitted models; stochastic heads represent within-model variability. Neither guarantees that all models will disagree on an unfamiliar case. Out-of-distribution (OOD) detection asks whether current inputs or transitions differ from the training conditions, not whether a prediction is certainly wrong.

Recent horizon-calibrated world-model work explicitly trains temporal uncertainty. Its action-free video pretraining setting is relevant but different from torque-conditioned contact prediction.[34] Physical Firmware would need held-out coverage tests at multiple horizons, under contact transitions and under the planner's own selected actions. A variance forced to increase is not evidence of calibration, and increasing variance monotonically is not a universal physical law.

Technical note: the rollout distribution and calibration contract
\[p_\phi(z_{t+1:t+K},c_{t+1:t+K},\theta\mid b_t,a_{t:t+K-1},\theta_e).\]

Here \(K\) is the number of future transitions, \(b_t\) the current belief, \(\theta\) uncertain environment parameters assumed constant within this short rollout, and \(\theta_e\) embodiment parameters. The sequence has \(K\) actions, not \(K+1\). If environment parameters vary, predict their sequence as well. Integrate over uncertain embodiment parameters rather than treating their estimate as exact.

Rollouts should retain temporal correlations. Sample one model and a persistent parameter hypothesis for a trajectory where appropriate; independently resampling mass every step would represent a different physical process. Integrate mode probabilities with state uncertainty rather than attaching an unrelated confidence scalar after prediction.

Calibration means stated uncertainty matches empirical frequencies on a specified distribution. Check interval or region coverage together with sharpness: a region containing everything has coverage but little decision value. Conformal methods can calibrate sets under assumptions such as exchangeability; correlated robot trajectories and adaptive action selection violate simple exchangeability. Extensions exist, but arbitrary deployment shift is not covered for free.[33]

THE PLANNER CAN EXPLOIT THE MODEL

Optimization searches for low predicted cost. It can therefore select an action because the model is wrong there. Evaluate disagreement and error on optimized candidate actions, not only random held-out trajectories. Rejecting uncertain plans is useful only if uncertainty actually predicts their failures.

09 / TRAINING PROGRAMLearn broad structure. Then pay close attention to reality.

A plausible program combines simulation, synchronized real interaction, and deliberate embodiment variation. Simulation provides privileged state labels; real data exposes model mismatch. Cross-embodiment training tests whether robot identity has become an accidental shortcut. These stages can overlap; they describe responsibilities, not a proven recipe.

08 / Proposed training pipeline

  1. Structured simulation pretraining

    Varied bodies, contacts, materials, sensors, and action realization. Learn predictive structure with privileged labels.

  2. Multimodal real-world prediction

    Synchronize images, geometry, proprioception, force, touch, commands, and realized actuator state. Learn the discrepancy.

  3. Cross-embodiment training

    Share core operators; vary robot morphology and interfaces. Hold out whole embodiments.

  4. Deployment identification

    Freeze the core initially. Infer embodiment parameters; fit limited calibration and residual modules.

  5. Continual learning with regression gates

    Curate deployment residuals; replay prior tasks; verify constraints and retained performance before updates.

Figure 8. Simulation → real interaction → cross-embodiment learning → adaptation. Deployment feedback may inform later core releases, but evaluation episodes must not leak into training.

A / Simulation teaches a family of possible interactions

Randomize mass, inertia, friction, compliance, damping, geometry, object count, morphology, sensor noise, and action latency. Use rigid-body simulation first; add deformable or differentiable simulation when the task requires it. Ground-truth pose, force, mode, and parameter labels help separate failures of perception from failures of dynamics.

Differentiable simulators can support parameter fitting and gradient-based planning; DiffTaichi is one established route to differentiable physical computation.[32] Differentiability does not remove the reality gap. Simulation pretraining should teach variation and structure, not certify that contact labels, friction, or sensor models match hardware.

B / Real prediction needs synchronized interventions

Record RGB, depth, joint positions and velocities, measured or estimated torques, wrist force/torque, tactile signals where available, commanded actions, actuator state, timestamps, outcomes, and failures. Record units, frames, calibration versions, sampling intervals, and missing-modality masks. Command timestamps and observed execution timestamps are different data.

Predict future features, geometry, motion, contact events, and forces. Not all real data has all labels: mask unavailable terms, use sensor likelihoods, and distinguish estimated contact annotations from instrumented ground truth. Resample with care around impacts; interpolating a discontinuity can erase the event that matters.

Passive videos supply visual and temporal priors but generally do not identify the effect of a robot command. Observed actions can be correlated with hidden causes. Controlled interventions, varied action coverage, and explicit action realization are needed before interpreting a predictor causally. A model that imitates the dataset's typical future is not necessarily accurate for a planner's counterfactual action.

C / Reuse must be exercised during training

Train with several morphologies, grippers, camera configurations, and control interfaces where feasible. Normalize units and frames without erasing meaningful differences. Test held-out combinations and a wholly held-out robot, not merely a new episode from a familiar robot. Public cross-embodiment datasets are valuable, but their sensor and force coverage must be audited for this objective.[7]

D–E / Adaptation and continual learning are separate budgets

At deployment, first update embodiment parameters, small residuals, and calibration layers. If that fails, allow a separately reported core-update condition. Continual learning can later incorporate new residuals and failure cases using replay, constrained parameterizations, and a fixed regression suite. Catastrophic forgetting—losing previous capabilities while learning new ones—must be measured across old robots and tasks. A held-out benchmark is not a replay buffer.

10 / TRAINING OBJECTIVESOptimize the predictions the controller will actually use.

A model can predict the next frame well and fail after repeated rollout. Training should expose it to its own predicted states, to contact transitions, and to several horizons. Latent overshooting compares multi-step predicted latent distributions with later inferred states; it was used in PlaNet.[19] Scheduled rollout training gradually replaces ground-truth context with predicted context. Neither automatically solves the difference between training data and a planner's chosen actions.

Technical note: one candidate multi-objective loss
\[\begin{aligned}\mathcal L={}&\lambda_{\rm roll}\mathcal L_{\rm roll}+\lambda_c\mathcal L_c+\lambda_f\mathcal L_f\\&+\lambda_g\mathcal L_g+\lambda_e\mathcal L_e+\lambda_s\mathcal L_s\\&+\lambda_u\mathcal L_u+\lambda_{\rm cal}\mathcal L_{\rm cal}+\lambda_r\mathcal L_r.\end{aligned}\]

Each nonnegative \(\lambda\) weights a term after units and scales are normalized. The terms below are a design menu, not nine objectives known to be jointly optimal.

TermPurpose and limitation
Rollout, LrollMulti-horizon likelihood or proper predictive score for future observable targets; not just next-state error.
Contact, LcMode/event likelihood and timing; handle imbalance without destroying probability calibration.
Force, LfForce/torque or impulse prediction with sensor noise and bandwidth modeled.
Geometry, LgPose, distance, shape, identity, and geometric consistency for available labels.
Equivariance, LeConsistency under valid transformations when symmetry is not enforced exactly.
Structure, LsPenalty for unenforced constraints, inappropriate energy gain, or invalid contact. Avoid duplicate constraints.
Uncertainty, LuProper distributional scores for additional probabilistic heads; omit duplicate NLL already in rollout loss.
Calibration, LcalValidation-tuned calibration or coverage surrogate, with a separate untouched evaluation set.
Representation, LrTemporal, cross-modal, or latent consistency without collapsed representations.
\[\mathcal L_{\rm roll}=-\mathbb E_{\mathcal D}\!\left[\sum_{k=1}^{K}w_k\log p_\phi(y_{t+k}\mid b_t,a_{t:t+k-1})\right].\]

Here \(\mathcal D\) is the training trajectory distribution, \(K\) the maximum rollout horizon, \(w_k\geq0\) horizon weights, and \(y\) an observed target such as pose, geometry, or sensor features. This is a sum of marginal predictive negative log likelihoods, not the joint likelihood of the entire trajectory. Missing labels require masks. The distribution must arise from the same rollout and observation interface used during planning.

A learned latent target needs an anchored decoder or a noncollapse objective; reducing latent error by shrinking all features is not learning physics. A likelihood must specify noise and units. If the predictor only produces samples, use an appropriate sample-based proper score instead of claiming a tractable NLL. A hard architectural constraint needs no penalty to “enforce” the same property. Measure whether each remaining loss improves downstream decisions.

Use rollout corruption and diverse action sequences to expose compounding error. Collect new interactions where the current planner fails, with a separate safety and data-governance process. Evaluate on independently collected episodes and splits by object, parameter combination, and embodiment. Randomly splitting adjacent frames would make the apparent generalization nearly meaningless.

11 / SYMMETRYTransform the problem consistently.

Equivariance means that transforming an input transforms the output in the corresponding way. A vector force should rotate when the coordinate frame rotates; a scalar mass should not. Equivariant graph networks provide tools for building such behavior into learned computations.[28]

The world is not indiscriminately SE(3)-symmetric. SE(3) is the group of three-dimensional rotations and translations. Rotating a scene while keeping gravity fixed can change what happens. A fixed robot base, asymmetric gripper, fixture, camera, and external field also matter. A change of coordinates transforms those quantities too; physically rotating only the object is a different intervention.

Technical note: gravity-conditioned equivariance
\[F(g\!\cdot\!x,\ g\!\cdot\!u,\ g\!\cdot\!\gamma)=g\!\cdot\!F(x,u,\gamma).\]

Here \(F\) predicts a physical state or compatible vector field, \(g\) is a valid rigid transformation, and \(\gamma\) collects gravity and other environmental context. The dot denotes the appropriate group action on each quantity: vectors rotate, points rotate and translate, and invariant scalars stay fixed. Actions must transform according to their meaning; joint commands do not transform like Cartesian forces.

For a physical transformation that keeps \(\gamma\) fixed, only transformations preserving that context are symmetries. Observation encoders have additional visibility and camera constraints. Apply equivariance to selected geometric components and test with transformed gravity, robot, and fixture descriptions where required. A generic SE(3) augmentation of images is not the same guarantee.

The expected benefit is fewer examples needed to learn the same relation in different frames. The risk is forbidding a real asymmetry. Compare architectural equivariance with valid data augmentation and with an equally expressive unconstrained model.

12 / NUMERICAL INTEGRATIONThe integrator is part of the model.

A continuous-time dynamics equation is not a trajectory. A numerical integrator advances it in finite steps. Explicit Euler is simple but can accumulate severe error. Runge–Kutta methods improve local accuracy for smooth dynamics. Symplectic and variational methods preserve selected geometric structure; for appropriate smooth Hamiltonian problems, their long-time behavior can be much better than a generic discretization.[30]

They do not universally conserve exact energy, resolve stiff contact, or guarantee accurate control. Near an impact, event detection and reset handling may matter more than smooth-step order. Persistent constraints can require stabilization or projection. Stiff compliant contact may need implicit updates or small time steps. Adaptive stepping changes computational cost and can interfere with geometric guarantees unless designed for them.

Proposed design: train and evaluate the actual discrete rollout used at deployment, including its contact solver, tolerances, and time step. Compare equal wall-clock budgets as well as equal step counts. A model trained with one solver may compensate for that solver's error and fail when integrated differently. Report both physical constraint violations and task outcomes.

13 / PHYSICAL FIRMWARE + VLAA consequence model can complement a generalist policy.

The useful shorthand is: VLAs learn what action to take; world models predict what may happen; simulators encode an explicit approximation of what should happen. Physical Firmware asks whether selected architectural physics improves a reusable learned predictor. These are roles, not mutually exclusive model species.

Flow and diffusion policies learn distributions of action sequences, which is valuable when several behaviors can satisfy the same instruction. Diffusion Policy and π0 establish two important action-generation lineages.[8][1] A distribution over good actions is different from a distribution over the consequences of an arbitrary proposed action. A single architecture can learn both.

07 / Five possible integration patterns

  1. A / Direct policy

    Observation + instruction → VLA → actions

    The strong baseline: reactive or history-conditioned behavior, without an exposed rollout service.
  2. B / Consequence critic

    VLA proposals → firmware predictions → select / modify → controller

    Compare candidate consequences. Rejection is only as good as model calibration and coverage.
  3. C / Semantic goal + MPC

    VLA subgoal → cost / constraints → firmware + action search → controller

    A model-predictive controller searches locally; translating language into a valid objective is still a separate problem.
  4. D / Training infrastructure

    Firmware → counterfactuals / hard cases → policy training → fast VLA

    Use predicted experience offline, with real validation and controls on model bias.
  5. E / Hybrid runtime

    VLA sequencing ⇄ local physical planner ⇄ stabilizing controller

    Separate time scales; use measured feedback to correct all three levels.
Figure 7. Alternative architectures, not a maturity ranking. The simplest effective integration should win. None is a safety certification.

Predictive auxiliary objectives already appear in policy learning. GR00T N1.5 adds Future LAtent Representation Alignment (FLARE) to action learning, aligning representations with future embeddings rather than generating future frames.[10] π0.7 uses richer conditioning, including visual subgoals produced by a lightweight world model.[4] These are concrete evidence that prediction and action learning are converging.

The distinction to test is therefore sharper than “has a world model.” Does an explicit, reusable physical rollout interface—with identified action realization, contact structure, and calibrated uncertainty—add value beyond the predictive representations already inside a policy? The answer could be no. An auxiliary loss might capture the useful structure at lower runtime cost.

14 / PLANNINGPredict. Choose. Act briefly. Observe again.

Model-predictive control (MPC) repeatedly evaluates action sequences over a finite horizon, executes the first action or short prefix, updates its state estimate, and replans. A sampled optimizer can compare future outcomes without differentiating through every contact. Gradient-based planning is another option when the model and its derivatives are reliable.

Cost can include distance to a goal, time, energy, force, constraint violation, and task failure. Risk-sensitive planning accounts for the distribution of outcomes rather than just its mean. A rare drop or damaging contact may deserve more weight than a small improvement in average speed. Uncertainty penalties should be validated for usefulness, not selected because cautious trajectories look reassuring.

Technical note: risk-aware MPC and its limits
\[\mathbf a^*=\arg\min_{\mathbf a\in\mathcal A_K}\left\{\mathbb E_{\tau\sim p_\phi(\cdot\mid b_t,\mathbf a)}[C(\tau)]+\beta U(\mathbf a,b_t)\right\}.\]

The candidate sequence \(\mathbf a=(a_t,\ldots,a_{t+K-1})\) belongs to \(\mathcal A_K\), the set respecting command and actuator limits. The random trajectory \(\tau\) comes from the rollout distribution; \(C\) is accumulated task cost; \(U\) is a chosen uncertainty penalty; and \(\beta\geq0\) sets its weight in compatible units. This objective is a proposed design, not a theorem that uncertainty penalties produce safety.

\[\Pr_{p_\phi}\{\tau\notin\mathcal S\}\leq\varepsilon.\]

A model-based chance constraint can bound predicted probability of leaving a chosen admissible set \(\mathcal S\), at tolerance \(\varepsilon\). The real-world bound is only as valid as the model and calibration assumptions. Tail-cost objectives such as conditional value at risk emphasize bad outcomes; they also require enough samples of those outcomes to be meaningful. Independently enforced robot limits and an explicit fallback remain necessary.

The open-loop candidate sequence does not include every possible future observation and recovery action. Replanning partly compensates, but full belief-space or dual-control planning—which chooses actions both to control and to learn—costs more. Do not count the benefit of future feedback twice in an open-loop rollout.

Latency is an architectural constraint

High-frequency motor stabilization belongs in a suitable controller. Medium-horizon physical planning can run more slowly; semantic sequencing can run more slowly still. Their exact frequencies depend on the robot, task, sensors, and communication path. A predictor taking seconds cannot sit synchronously inside a loop requiring a new decision every 100 milliseconds; this is a timing example, not a measured system specification.

Measure end-to-end deadlines, including state estimation, candidate generation, integration, uncertainty sampling, and transport—not just one neural forward pass. Warm-start planning, batch candidates, shorten horizons, distill policies, and use hierarchical models when useful. If a deadline is missed, use a defined safe hold or fallback appropriate to the hardware. Model-based RL also teaches that short, trusted rollouts may outperform extensive use of a biased model.[21]

Proposed interface: the data contract a planner would need
# Interface sketch — no implemented library is claimed.
belief = estimator.update(history, robot_description)
grounding = adapter.infer(calibration_history, belief)
prediction = core.rollout(
    belief=belief,
    command_sequences=candidates,
    embodiment=grounding,
    sample_times=planning_times,
    deadline=remaining_planning_budget,
)
# Return correlated trajectory samples, contact modes,
# forces, parameter hypotheses, validity diagnostics,
# reference frames, units, timestamps, and model version.
plan = planner.evaluate(prediction, goal, constraints)
controller.execute_first_prefix(plan)
# Compare actual consequences; update belief and replan.

A prediction response should specify whether its uncertainty is calibrated on comparable conditions, which assumptions failed, and whether computation met its deadline. An invalid prediction is a possible output, not an exception to hide.

15 / WORLD MODEL VS. SIMULATORThe useful distinction is what each model commits to.

A classical simulator specifies geometry, equations, parameters, and a numerical approximation. Its state is inspectable and interventions explicit. Its weaknesses are imperfect identification, simplified contact and materials, and the simulation-to-reality gap. It can already incorporate learned parameters or residuals.

A generative video world model predicts rich visual futures. Broad visual training can cover scenes and behaviors difficult to author manually. But visually plausible motion is not necessarily accurate under a particular actuator command, and a convincing image does not establish correct force or contact. Action-conditioned variants need a precise account of what the action channel means.

A latent dynamics model predicts compact hidden representations. It can be efficient for planning without rendering images, but its state may be hard to inspect or constrain. A differentiable simulator provides derivatives through its approximate dynamics, enabling identification or optimization; differentiation is a computational capability, not a guarantee of realism.

Physical Firmware belongs in this overlapping space. It proposes learned dynamics plus selected physical structure, real-world residuals, explicit uncertainty, and a reusable embodiment interface. If implemented as a structured learned simulator, that description is entirely appropriate. It should earn its separate name through useful transfer, not through a claim to a previously empty category.

  • Physical consistency ≠ prediction accuracy.
  • Prediction accuracy ≠ planning usefulness.
  • Visual realism ≠ physical accuracy.
  • One-step accuracy ≠ stable rollout ≠ successful control.
  • Simulation transfer ≠ real-world transfer.
  • Architectural prior ≠ scientific truth.

16 / WHAT IS ACTUALLY UNIVERSAL?Universal laws do not imply a universal learned representation.

A rigid-body state can transfer across many objects while failing on cloth. A contact operator useful for a parallel gripper may miss the distributed compliance of a hand. Tabletop manipulation and locomotion share mechanics but differ in reachable states, contact topology, sensing, and the consequences of failure.

The transferable unit might be mechanics, a set of contact primitives, a geometric representation, a parameter-inference procedure, or simply broad pretraining. Those alternatives make different predictions. If a pretrained unstructured model adapts equally well, useful reuse exists, but the claim for explicit architectural physics weakens.

Speculation: the eventual design may be a hierarchy or a learned mixture of specialized models: rigid-body, articulated-body, deformable-object, fluid-interaction, and tool-interaction modules. Such a mixture needs a routing rule, compatible state and force interfaces, and uncertainty about choosing the wrong model. Joining individually plausible modules does not guarantee that their combined energy and contact behavior is consistent.

Start with a declared domain. “Reusable within contact-rich rigid tabletop manipulation” is a meaningful claim. “Universal physical intelligence” is not a justified extrapolation from it. Cross-object transfer, cross-task transfer, and cross-embodiment transfer must be reported separately.

17 / THE SEPTEMBER 2026 LANDSCAPEThe categories are converging.

Physical Intelligence's progression matters because it strengthens the behavioral baseline. π0 combines broad robot data with a flow-based action model; π0.5 addresses open-world generalization; π*0.6 uses learning from experience and reinforcement-learning post-training; π0.7 adds steerability and reports compositional and cross-embodiment behavior.[1][2][3][4] These results concern the authors' evaluated settings; they are not a claim of arbitrary task competence.

Its 2026 Multi-Scale Embodied Memory work adds long- and short-term history, while RL Token work uses compact VLA representations for online policy improvement.[5][6] Memory matters because observation history can disambiguate state. Online adaptation matters because frozen imitation is no longer the only policy baseline. Neither, on its own, implies an explicit posterior over physical parameters.

NVIDIA's GR00T N1 is a generalist humanoid policy; N1.5 couples policy training to future representation alignment. Later N1.6 and N1.7 releases extend the available model family; the official repository identifies N1.7 as its current release and describes relative end-effector actions shared across human and robot data.[9][10] Action representation itself can support transfer. The other important precedent is joint prediction and action learning. An auxiliary embedding objective should not be mistaken for a fully identified, externally queryable contact simulator.

NVIDIA Cosmos develops world foundation models for physical AI. The original platform emphasizes large-scale world modeling; Cosmos 3's 2026 technical report describes an omnimodal model spanning understanding, generation, action, and forward/inverse dynamics.[11] This makes a strict “video model versus policy” taxonomy increasingly inadequate. Capabilities described in a technical report and released model availability should still be distinguished.

Google DeepMind's Gemini Robotics work joins embodied reasoning to action, while Gemini Robotics ER 2, announced in July 2026, emphasizes video-based progress monitoring, task orchestration, and delegation to lower-level execution systems.[12] Genie 3 demonstrates interactive generated environments; such environments are relevant to experience generation, but visual interactivity alone does not establish calibrated robot-contact dynamics.[13]

Two 2026 surveys help organize this convergence: Wang and colleagues distinguish representations and how prediction is connected to action; Kirchner, Purschke, and Knoll foreground uncertainty and closed-loop control.[14][15] The comparison below is an architectural map, not a benchmark ranking. Entries marked “varies” depend on implementation and training data.

Approach / examplesOutputPhysical structureAction conditioningCross-embodimentContactUncertaintyPrimary role
VLA / π family; Gemini RoboticsActions / action sequencesMostly learned; architecture variesGenerates actions from contextTrained / evaluated in multiple settingsOften implicit; sensor-dependentAction diversity is not calibrated outcome uncertaintyGeneralist behavior
VLA + predictive objective / GR00T N1.5Actions + future representation supervisionLearned visual/action representationsJoint training; exposed rollout API not impliedMulti-embodiment dataImplicit in described objectiveNot certified by auxiliary predictionRepresentation and policy learning
Generative world foundation model / Cosmos; Genie 3Video / multimodal futures; some models also actionsLearned; geometric conditioning variesVaries: control, video, or robot-action channelsRequires action and embodiment groundingVisual plausibility is insufficientSample diversity; calibration must be testedGeneration, predictive infrastructure
Classical physics simulatorPhysical trajectoriesExplicit equations and geometryExplicit controls / forcesNew robot models and parametersChosen numerical contact lawUsually added through parameter/noise modelsSimulation, control, synthetic data
Differentiable simulator / DiffTaichiTrajectories and derivativesExplicit or hybridExplicit, model-dependentRe-identification / remodelingGradients depend on formulationNot automaticIdentification, optimization
Learned latent dynamics / PlaNet; PETSLatent or state futuresLearned transitions; priors varyExplicit candidate actionsNot automaticOften learned implicitlyStochastic state / ensemblesPlanning and model-based RL
Physics-structured dynamics / HNN; port-Hamiltonian modelsState derivatives / rolloutsSelected energy / geometry constraintsDepends on model; ports in controlled variantsMust be demonstratedSeparate extension usually neededDepends on implementationDynamics learning, control
Physical Firmware / proposalGrounded probabilistic physical futuresSelected graph, geometry, dynamics, contact priorsCommands through identified realizationCentral hypothesisExplicit modes / constraints + learned correctionsBelief, parameters, modes, model uncertaintyReusable prediction layer

On narrow screens, scroll the comparison horizontally. “Not automatic” is not a claim that a capability is impossible or absent from every system in that category.

18 / THE DECISIVE EXPERIMENTMeasure the cost of reaching the same reliability.

The central prediction is a leftward shift: less new real-world experience to reach the same reliability after a physical change.

Start with contact-rich tabletop manipulation: pushing, repositioning, constrained sliding, stacking, tray placement, pick/place under mass variation, and later peg or connector insertion. Hold task semantics approximately constant while changing physical conditions. A better nominal score is useful, but it does not by itself demonstrate reusable physical knowledge.

09 / The claim to test — conceptual curves, not results

Conceptual data-efficiency hypothesisThree dashed illustrative curves show reliability increasing with new real-world task-specific interaction. The proposed structured model reaches a target earlier than two illustrative baselines. The ordering and gap are unknown, with no measured values shown. Predeclared reliability target Possible data saving HYPOTHESIS · NO EMPIRICAL VALUESTask reliabilityNew real-world task-specific interaction →Include calibration and failures in the cost ledger Hypothetical reliability curves, mobile diagramThree dashed curves illustrate a possible leftward shift for the structured candidate, reaching the same reliability with less new interaction. No results or numerical improvement are claimed. HYPOTHESIS · NOT DATA TargetReliabilityNew real interaction →Include calibration + failures
Structured candidatePretrained policy baselineUnstructured model baseline
Figure 9. An illustration of the desired comparison, not a forecast of the ordering. Curves may overlap or reverse. No numerical improvement is assumed.

Baselines that could actually disprove the claim

ArmSystemWhat the comparison asks
ATask policy trained from scratchWhat is the total gain over task-only learning?
BPretrained VLA / behavioral policy, fine-tuned; include a viable online-RL variantDoes it beat strong behavioral reuse and adaptation?
CUnstructured learned world model + the same plannerDoes architectural structure add anything beyond predictive pretraining?
DSimulator-trained policy; also a calibrated simulator + MPC where feasibleIs learning the proposed core better than existing simulation and identification?
EEstablished physics-informed / structured dynamics modelDoes the proposed combination add value over known structure?
FPhysical Firmware candidate + MPCDoes its prediction layer reduce adaptation effort?
GPhysical Firmware representation + learned task policyDoes reuse help without expensive runtime search?
HAblations: no equivariance, contact structure, energy structure, uncertainty, or shared coreWhich component causes any improvement?

Hold observation quality, action interface, demonstration quality, task data, simulator access, planner budget, and evaluation episodes constant wherever meaningful. Match capacity and training compute for controlled model comparisons. A large pretrained VLA cannot always be matched exactly; report both a controlled scientific comparison and a practical best-available-system comparison, including inherited data and compute.

Plot reliability against real interaction time and episode count at several adaptation budgets. Record all failed attempts and calibration, not only demonstrations used for gradient updates. Report upstream pretraining separately and show amortization across deployments; an expensive reusable core has not saved data in total merely because downstream fine-tuning is small.

Pre-register the reliability target, minimum practically meaningful saving, success tolerance, action deadline, exclusions, and stopping rule. Use paired physical conditions, multiple training seeds, and uncertainty intervals clustered by object or session. Choose the number of trials with a power analysis based on pilot variability. If a method never reaches the target, report that rather than extrapolating its curve.

19 / DISTRIBUTION-SHIFT TESTSChange what the robot sees separately from what the world does.

A distribution shift is a change between training and deployment conditions. A new camera angle tests perception. A hidden mass change tests dynamics inference. Combining them immediately makes it difficult to identify the reason for success or failure.

Perceptual shift

Camera pose, lighting, background, visual texture, depth noise, occlusion, and sensor dropout. Hold physical parameters fixed when possible.

Physical shift

Mass, friction, geometry, object size, center of mass, compliance, and support configuration. Keep visual appearance controlled where possible.

Compositional shift

Object count, clutter, starting pose, combinations of familiar objects, contact topology, and changed fixtures. Some shifts affect both perception and dynamics.

Embodiment / interface shift

New arm, gripper, controller, latency, control frequency, or sensor suite. Report these separately from changing an object on the same robot.

Distinguish interpolation within trained parameter ranges, held-out combinations, and extrapolation beyond those ranges. Use one-factor tests to diagnose, then joint shifts to test practical robustness. Oracle state and oracle parameter conditions reveal how much failure comes from estimation rather than dynamics. Finally test the full sensor-to-action system: privileged-state results alone do not establish deployment transfer.

20 / METRICSPrediction, control, adaptation, and operating cost.

Success rate is necessary but insufficient. A model might reach the goal by using damaging forces, excessive resets, or slow planning. Conversely, lower prediction error on irrelevant pixels may improve none of the outcomes that matter.

LevelReportInterpretation guardrail
PredictiveOne- and multi-step pose/velocity error; rollout NLL where defined; force error; contact event precision, recall, F1 and timing; geometry consistency; constraint violation; applicable energy/dissipation residualsDisaggregate by contact mode and horizon. Define units, event matching windows, and reference sensors.
UncertaintyCoverage and sharpness; reliability curves; NLL; Brier score for event probabilities; OOD response; calibration by horizon and modeScore optimized actions too. Broad intervals and large variance are not sufficient evidence of useful calibration.
ControlSuccess, recovery, unsafe failure and intervention rates; force-limit violations; model exploitation; latency distribution; deadline misses; action frequencySeparate model-predicted safety from actual outcomes. Average latency hides missed deadlines.
AdaptationDemonstrations; all interaction time; calibration episodes; wall-clock time; gradient updates; fraction and number of parameters changed; retention on prior robotsDo not classify calibration as “free” or cross-object transfer as cross-embodiment transfer.
OperationalEngineering and operator hours; deployment time per new variant; downtime; reset labor; compute cost; hardware/failure costA data saving can be outweighed by integration or maintenance overhead.

Model exploitation should have a concrete diagnostic: compare predicted and realized cost for actions selected by progressively stronger search, within the same bounded action space. If more optimization improves predicted cost while worsening measured outcomes, the planner is finding model defects. Log those trajectories as failures of the prediction–planning system.

Contact F1 and Brier scores answer different questions: event detection quality and probability quality. NLL is meaningful only for a declared predictive density. Energy consistency applies only to modeled components with a meaningful energy balance. None should be used as a universal substitute for control evaluation.

21 / FALSIFICATIONWhat would prove the thesis wrong?

A broad philosophical claim about all possible architectures cannot be refuted by one failed implementation. A scoped engineering hypothesis can. For the declared task family, data budget, and model class, specify an effect large enough to justify the added structure, then give the competing systems a fair chance to match it.

ABANDON OR NARROW THE PROPOSAL IF…
  • A matched unstructured world model transfers equally well or better.
  • The same reliability needs approximately the same amount of task-specific real interaction once all adaptation is counted.
  • Adapter calibration costs as much as policy adaptation or predictor retraining.
  • Physical priors systematically exclude real behavior, and residual capacity removes any useful distinction from an unstructured model.
  • Uncertainty remains miscalibrated where the planner needs it, or conservative penalties erase the claimed efficiency gain.
  • Contact prediction fails in held-out impact, slip, or multi-contact regimes.
  • The planner repeatedly exploits model errors despite reasonable uncertainty and rollout controls.
  • Cross-embodiment retention is negligible after a fair grounding procedure.
  • Long-horizon stability improves but successful control and adaptation cost do not.
  • VLA pretraining or predictive auxiliary objectives already obtain equivalent reusable physical knowledge at lower total cost.

An underpowered null result is inconclusive. A repeatable equivalence result with uncertainty intervals narrow enough to exclude the predeclared useful improvement is damaging evidence. If only one component helps—say, contact geometry—retain that result and discard the unsupported claim for the larger architecture. If the best solution is a conventional identified simulator, use it.

Conversely, one successful pushing experiment would support a limited candidate, not universality. The stronger claim needs transfer across physical changes and embodiments, reliable uncertainty under planning, and a measurable reduction in total adaptation effort.

22 / RESEARCH ROADMAPMake each stage earn the next one.

10 / Proposed program, with decision gates

  1. P0
    Mathematical prototype

    Pendulum, cart pole, spring–mass, articulated chain. Verify parameterization, energy accounting, and integration. Gate: correct constrained behavior without hidden numerical artifacts.

  2. P1
    Rigid-object interaction in simulation

    Pushing, sliding, collisions, variable mass and friction. Gate: useful prediction on held-out combinations and contact graphs.

  3. P2
    One real tabletop system

    Arm, RGB-D, wrist force/torque. Gate: measured contact prediction and lower adaptation cost after identification.

  4. P3
    Tactile contact

    Insertion and constrained manipulation. Gate: compatible tactile representation improves closed-loop outcomes over matched vision/force baselines.

  5. P4
    Second embodiment

    Another arm or gripper with the core frozen. Gate: calibration beats retraining in total cost while retaining earlier performance.

  6. P5
    VLA integration

    Generalist policy for semantics and proposals, firmware for consequences. Gate: incremental benefit over the same VLA with equal sensing and compute.

  7. P6
    Industrial workflow

    Repeated manipulation with changing parts. Gate: lower deployment cost at the required reliability across operating shifts.

Figure 10. A proposed research sequence, not a delivery schedule. Failure at a gate should revise the model or stop expansion.

P0 should be deliberately small. A successful pendulum validates code and assumptions; it cannot justify a claim about contact-rich robotics. The scientific center of gravity is P2–P4: real contact, identifiable grounding, and a core that survives a changed body.

23 / THE FIRST REAL EXPERIMENTOne arm, variable objects, and an honest comparison.

Proposed minimum credible experiment: use a Franka Panda or equivalent research arm with joint-state access, an RGB-D camera, and a calibrated wrist force/torque sensor. Begin with planar pushing and object repositioning on an instrumented, bounded tabletop. Tactile sensing and insertion are follow-ups, not prerequisites for the first result.

Make hidden physical changes measurable

Use objects with interchangeable internal weights and adjustable center of mass, several geometries and sizes, and replaceable surface materials. Measure reference mass, geometry, and friction under a declared procedure for evaluation; do not feed held-out parameter labels to the deployed model. Keep some appearances constant while changing weight so a visual shortcut cannot explain transfer.

Create disjoint splits for familiar objects, held-out combinations of known factors, and new parameter ranges. Reserve entire sessions and object configurations for evaluation. Use synchronized pose tracking as an evaluation reference—fiducials or another calibrated tracker are acceptable for the prototype—and report its error. Keep an oracle-state condition to diagnose perception, but make the main result use the same available sensors for both models.

Build two predictors that differ in the thing being tested

The unstructured candidate uses the same state estimator, object representation, probabilistic output family, training trajectories, and planner. It predicts transitions without the selected mechanical constraints. The structured candidate uses the same inputs with explicit geometry, an appropriate rigid-body update, a contact treatment, and a limited learned residual. Match capacity and compute as closely as possible. Then add ablations for graph structure, equivariance, and uncertainty separately.

For slow planar pushing, a quasi-static model—one that neglects inertia when its effects are small—may be the strongest structured candidate. Include it. Introduce faster motion and changes of velocity only if testing inertial structure. Otherwise a Hamiltonian network could appear ineffective simply because the experiment never required its contribution.

Implementation note: a minimal candidate worth building

Use explicit planar pose and velocity for each object; a robot tool node; shape features; a belief over friction and inertial properties; and actuator history. A recurrent encoder integrates visual pose estimates, joint states, and forces. A probabilistic transition ensemble predicts short action-conditioned trajectories, including contact mode and sensor-level force targets.

For the structured variant, separate free motion, contact response, and bounded residual prediction. Use an analytic or constrained contact update rather than assuming smooth energy dynamics will discover impacts. Estimate common physical parameters from a short calibration history. Preserve the same observation decoder and uncertainty family in the unstructured control.

Use a sampled MPC optimizer such as the cross-entropy method: repeatedly sample candidate command sequences, keep lower-cost candidates, and refit the sampling distribution. Give both predictors the same candidate budget and deadline. A separately tuned planner condition can test best practical performance, but must not replace the matched-planner comparison.

Run an adaptation ladder, then change the body

Evaluate before task-specific adaptation and after predeclared cumulative interaction budgets. Freeze the pretrained models between budget checkpoints. Count probing, failures, demonstrations, and autonomous exploration; report training time separately. Measure multi-step motion and force error, contact timing, parameter-shift transfer, calibration, MPC success, and the interaction needed to reach the same goal tolerance and reliability.

Next, change the arm or gripper while retaining the trained physical core. Allow only embodiment identification and limited adapter updates in the primary transfer condition. Compare against a freshly fitted predictor, a fine-tuned core, and the same policy baseline. A new gripper tests end-effector transfer; a genuinely different arm and control interface is a stronger embodiment test. Label each accurately.

THE MINIMUM IMPORTANT OUTCOME

A repeatable reduction in total new physical interaction at matched reliability, on held-out dynamics and then a changed embodiment, with the improvement traceable to specific structure. No fixed multiplicative gain is assumed.

24 / PRODUCT IMPLICATIONSOnly after the transfer result, ask where it pays.

Conditional product hypothesis: high-mix industrial manipulation could be a useful first application. The robot repeats a family of tasks while parts, fixtures, weights, or surface conditions change. A known workspace, measurable contact, robot telemetry, and explicit reliability targets make evaluation more tractable than a general household robot.

The value would be fewer engineering and operator hours to deploy a new variant, less downtime during changeover, or fewer expensive contact failures. These are testable operational outcomes. They must exceed the cost of extra sensors, calibration, inference hardware, integration, and maintaining the model.

A useful prediction layer might become a calibrated rollout service, an adaptation toolkit, or offline training infrastructure. The experiment should decide which form is justified. There is no market-size estimate or assumption that a technically elegant model becomes a standalone business.

25 / THE DATA FLYWHEELThe scarce asset would be grounded interaction history.

If the hypothesis works, repeated deployments could accumulate synchronized physical interaction data, rare failures, contact transitions, unusual materials, embodiment adapters, calibrated robot models, learned residuals, and cross-robot transfer records. A benchmark suite and a reliable account of deployment costs would be part of that asset.

The feedback loop is concrete: log prediction errors → identify a failure regime → collect bounded informative interactions → update a candidate → retest old and new conditions → deploy only after regression checks. More data is useful when it covers missing interactions and retains provenance; repeated nominal successes can reinforce the same blind spots.

Any defensibility would come from the combination of architecture, interaction data, embodiment grounding, evaluation, and deployment feedback. “We use Hamiltonian networks” is not a defensible claim to the underlying idea. Data rights, instrumentation quality, and reproducibility matter as much as collection volume.

26 / BOUNDARIESWhat Physical Firmware is not.

  • A universal physics oracle, or a claim that Newtonian mechanics solves robotics.
  • A conventional rigid-body simulator with a new name, though a candidate may include one.
  • A replacement for a VLA, perception system, robot operating system, or low-level motor controller.
  • Proof that architectural physics beats scaling or broad behavioral pretraining.
  • A guarantee of cross-embodiment intelligence.
  • A claim that all relevant manipulation can be expressed through Hamiltonian mechanics.
  • An experimentally validated system or product.

World model ≠ policy. VLA ≠ simulator. System identification ≠ task learning. Uncertainty estimate ≠ calibrated uncertainty. Same task ≠ same dynamics. Keeping those distinctions visible is what makes the proposal testable.

27 / THE FINAL BETHow much knowledge can survive a change?

Robot learning is demonstrating that scale, heterogeneous data, and foundation models can produce increasingly general behavior. Physical Firmware asks a narrower question:

If a robot's predictive model is explicitly structured around geometry, constraints, contacts, uncertainty, and dynamical regularities, can more of what it learns survive the transition from one task, object, or robot to another?

The answer is not known. It is experimentally testable. That is the project.

Physical Firmware is a reusable, probabilistic, contact-aware and physically structured world model intended to reduce the amount of new physical experience required when a robot encounters a new task, environment or embodiment.

SOURCES / VERIFIED THROUGH 19 SEPTEMBER 2026References & further reading

Research lineage, not evidence that this proposal works. Peer-reviewed papers, preprints, and official reports are labeled separately. Dates refer to the cited publication or report, not a search engine's crawl date. No external benchmark numbers are reproduced.

  1. Physical Intelligence / Black et al. π0: A Vision-Language-Action Flow Model for General Robot Control.2024 · arXiv technical report, 2410.24164 · DOI. Flow-based action generation and multi-robot pretraining.
  2. Physical Intelligence. π0.5: a Vision-Language-Action Model with Open-World Generalization.2025 · arXiv technical report, 2504.16054 · DOI. Generalization beyond familiar deployment environments.
  3. Physical Intelligence. π*0.6: a VLA That Learns From Experience.2025 · arXiv technical report, 2511.14759 · DOI. Reinforcement-learning post-training / RECAP; distinguished from the π0.6 base policy.
  4. Physical Intelligence. π0.7: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities.2026 · arXiv technical report, 2604.15483 · DOI · official explanation, 16 April. Steerability, composition, and generated visual subgoals.
  5. Physical Intelligence. VLAs with Long and Short-Term Memory.3 March 2026 · official research report. Multi-Scale Embodied Memory (MEM); temporal context for policies.
  6. Physical Intelligence. Precise Manipulation with Efficient Online RL.19 March 2026 · official research report. RL Token representation and online actor–critic adaptation.
  7. Open X-Embodiment Collaboration. Open X-Embodiment: Robotic Learning Datasets and RT-X Models.2023 preprint / ICRA 2024 · arXiv DOI · project. Heterogeneous robot data and cross-embodiment policies.
  8. Chi et al. Diffusion Policy: Visuomotor Policy Learning via Action Diffusion.RSS 2023 · DOI: 10.15607/RSS.2023.XIX.026. Diffusion over action sequences.
  9. NVIDIA et al. GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.2025 · arXiv technical report, 2503.14734 · DOI. Generalist humanoid policy architecture and training.
  10. NVIDIA Research. GR00T N1.5: An Improved Open Foundation Model for Generalist Humanoid Robots.2025 · official research report. FLARE future-representation objective. Later N1.6 / N1.7 release context: official repository and N1.7 FAQ, checked 19 September 2026. Repository documentation is not a peer-reviewed benchmark.
  11. NVIDIA. Cosmos World Foundation Model Platform for Physical AI and Cosmos 3: Omnimodal World Models for Physical AI.2025 and 2026 · arXiv technical reports · original DOI / Cosmos 3 DOI. Broad world modeling and convergence of generation, action, and prediction.
  12. Gemini Robotics Team / Google DeepMind. Gemini Robotics: Bringing AI into the Physical World; Hansen and Xu, Introducing Gemini Robotics ER 2.2025 technical report, arXiv DOI; 30 July 2026 official ER 2 announcement. Embodied reasoning and lower-level action execution are distinct roles.
  13. Google DeepMind. Genie 3: A new frontier for world models.2025 · official research announcement. Interactive generated environments; not a robot contact-accuracy certificate.
  14. Fangyuan Wang et al. World Models for Robotic Manipulation: A Survey.2026 · arXiv survey, submitted 27 May; SmartBot, first published 31 August; DOI: 10.1002/smb2.70053. Representation and decision-coupling taxonomy.
  15. Sven Kirchner, Nils Purschke, and Alois Knoll. A survey of world models for physical AI with uncertainty representation and control.3 September 2026 · Discover Artificial Intelligence 6, article 1037 · DOI: 10.1007/s44163-026-02122-1. Closed-loop control and uncertainty.
  16. Zhiyuan Zhang et al. ContactWorld: What Representations Matter in Vision-Tactile World Models for Contact-Rich Manipulation.2026 · arXiv preprint; v1 11 June, v2 8 September · DOI. Revised title and representation study; no peer-review status assumed.
  17. Caoliwen Wang et al. WorldContact: A Contact-Centric World Model for Scalable Robot Learning.17 September 2026 · arXiv preprint, 2609.19600 · DOI. Deformable-object dynamics and policy-training data generation; very recent evidence.
  18. Yunfan Lou et al. Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation.7 June 2026 · arXiv preprint, 2606.08737 · DOI. Joint action, visual, and tactile modeling.
  19. Danijar Hafner et al. Learning Latent Dynamics for Planning from Pixels.ICML 2019 · PMLR 97 · arXiv DOI. PlaNet, latent planning, and overshooting.
  20. Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine. Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models.NeurIPS 2018 · arXiv DOI. PETS, ensembles, and uncertainty propagation.
  21. Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine. When to Trust Your Model: Model-Based Policy Optimization.NeurIPS 2019 · arXiv DOI. Model bias and controlled rollout usage.
  22. Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, and David Duvenaud. Neural Ordinary Differential Equations.NeurIPS 2018 · arXiv DOI. Continuous-time learned dynamics.
  23. Samuel Greydanus, Misko Dzamba, and Jason Yosinski. Hamiltonian Neural Networks.NeurIPS 2019 · arXiv DOI. Conservative dynamics through a learned Hamiltonian.
  24. Miles Cranmer et al. Lagrangian Neural Networks.2020 · arXiv preprint, 2003.04630 · DOI. Learned Lagrangian formulation.
  25. Thai Duong, Abdullah Altawaitan, Jason Stanley, and Nikolay Atanasov. Port-Hamiltonian Neural ODE Networks on Lie Groups for Robot Dynamics Learning and Control.IEEE Transactions on Robotics, 2024 · DOI: 10.1109/TRO.2024.3428433. Geometry, dissipation, actuation, and control.
  26. Fabian J. Roth, Dominik K. Klein, Maximilian Kannapinn, Jan Peters, and Oliver Weeger. Stable Port-Hamiltonian Neural Networks.NeurIPS 2025 · DOI: 10.52202/085713-1693. Stability under specified architectural conditions.
  27. Alvaro Sanchez-Gonzalez et al. Learning to Simulate Complex Physics with Graph Networks.ICML 2020 · arXiv DOI. Shared graph computations for learned simulation.
  28. Víctor Garcia Satorras, Emiel Hoogeboom, and Max Welling. E(n) Equivariant Graph Neural Networks.ICML 2021 · PMLR 139 · arXiv DOI. Equivariant geometric computations.
  29. David Stewart and Jeffrey C. Trinkle. An Implicit Time-Stepping Scheme for Rigid Body Dynamics with Coulomb Friction.ICRA 2000 · DOI: 10.1109/ROBOT.2000.844054. Impulse-based time stepping and simultaneous contacts.
  30. Ernst Hairer, Christian Lubich, and Gerhard Wanner. Geometric Numerical Integration: Structure-Preserving Algorithms for Ordinary Differential Equations.Springer, second edition, 2006 · DOI: 10.1007/3-540-30666-8. Geometric integration and its assumptions.
  31. Ashish Kumar, Zipeng Fu, Deepak Pathak, and Jitendra Malik. RMA: Rapid Motor Adaptation for Legged Robots.RSS 2021 · arXiv DOI. History-based adaptation; locomotion evidence, not manipulation transfer evidence.
  32. Yuanming Hu et al. DiffTaichi: Differentiable Programming for Physical Simulation.ICLR 2020 · arXiv DOI. Differentiable simulation infrastructure.
  33. Rina Foygel Barber, Emmanuel J. Candès, Aaditya Ramdas, and Ryan J. Tibshirani. Conformal prediction beyond exchangeability.Annals of Statistics 51(2), 2023 · DOI: 10.1214/23-AOS2276. Coverage under departures from standard exchangeability assumptions.
  34. Shenghua Wan, Le Gan, and De-Chuan Zhan. Learning to Be Uncertain: Pre-training World Models with Horizon-Calibrated Uncertainty.ICLR 2026 · conference paper. Horizon-aware probabilistic video pretraining; applicability to robot-contact prediction remains to be tested.

This article defines a research hypothesis, an architectural design space, and an evaluation program. Diagrams are explanatory schematics. The learning curves are hypothetical. No implemented Physical Firmware library, trained model, commercial deployment, or measured performance advantage is asserted.