Skip to main content

Share story

AI Science

Researchers recover suppressed answers by altering an AI model's internal signals

Researchers testing Minerva-7B found that its answers often concealed distinctions represented inside the model. In a preprint, they report that the model responded identically to 63.7% of 124 prompt pairs designed to test professional risk, despite showing internal differences between the paired requests. The team also tested 25 facts presented under five linguistic framings. When false assumptions suppressed a correct answer, removing the internal signal associated with the falsehood restored that answer in 11 of 25 cases. The results show a possible way to test whether an AI system represents a fact it fails to state, but the intervention was demonstrated on one model.
Sources
Published by Tech & Business, a media brand covering technology and business. This story was sourced from cs.AI updates on arXiv.org and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
Products
Products

Bird.com raises $450 million in JPMorgan-led debt financing

Bird.com, a customer messaging company, raised $450 million in debt financing led by JPMorgan Chase & Co., Bloomberg reported. The company is seeking to return cash to its investors and employees. The transaction adds debt capital...

Capital Infrastructure
Capital Infrastructure

Hubble Network raises $200 million for satellite Bluetooth network

Satellite startup Hubble Network raised $200 million in a new funding round, Bloomberg reported, taking its valuation to $1.6 billion. The company is working on a network of spacecraft intended to provide global Bluetooth connecti...

Security Infrastructure
Security Infrastructure

CISA adds two exploited Check Point flaws to federal fix list

The U.S. Cybersecurity and Infrastructure Security Agency added two Check Point Security Gateway flaws to its list of known exploited vulnerabilities and told federal agencies to apply fixes or mitigations by September 25. Check P...

Security Infrastructure
Security Infrastructure

Analysis identifies two flaws behind exploited MikroTik router takeover chain

CERT Polska has identified the two RouterOS SSH flaws behind a previously reported attack that can give intruders administrative control of exposed MikroTik routers without completing authentication. One flaw lets a connection rea...