Nullify preprint reports selective LLM forgetting without weight updates
Researchers report in a preprint that Nullify, a method designed to suppress specific memorized information in large language models, matches or surpasses established forgetting baselines on TOFU and MUSE while preserving model ut...
NEMORA preprint reports long-range atomistic learning at hundreds of thousands of atoms
Researchers report in a preprint that NEMORA, a method for learning interactions between atoms over long distances, handles systems containing hundreds of thousands of atoms with computation time and memory that scale linearly wit...
Preprint reports schema-free generation of valid enterprise test data
Researchers report in an arXiv preprint that their Generalist Populator agent generated enterprise data with 100% constraint satisfaction and 0.88 average marginal fidelity across ten simulated environments without accessing datab...
Preprint reports task-completion gains from agent-generated interfaces
Researchers report in an arXiv preprint that training an agent to generate interactive interfaces improved a 4B model's Pass@3 task-completion score from 9.33% to 58.00%. Their GenUI-Harness pairs an agent that retrieves informati...
DreamTrue researchers report fewer interaction defects in robot video predictions
Researchers report that DreamTrue, a model that predicts videos of robot actions, reduced human-assessed interaction defects from 48.12% to 6.25% on AgiBot. The preprint addresses predictions that follow commands inaccurately or f...
FAITH preprint reports humanoid safety gains while preserving task performance
Researchers report that their FAITH safety filter achieved a 99.95% safety rate while retaining 97% of unfiltered task return on a 29-degree-of-freedom humanoid in Walking-Avoid. The preprint also describes demonstrations of the s...