Skip to main content

Share story

Science AI

Researchers Propose 'Psychological Competence' as Missing Dimension in AI Evaluation

Current AI evaluation frameworks focus on technical performance, accuracy, robustness, reasoning, and policy compliance, but overlook how human-facing AI systems affect user cognition, emotion, trust, and decision-making, researchers argue in a new position paper. As language models increasingly serve as advisors, coaches, tutors, and companions, their responses shape how users reason, interpret emotions, form beliefs, calibrate trust, and make decisions. The paper introduces "psychological competence" as the capacity of a human-facing AI system to support user cognition, emotional interpretation, and behavioral decision-making in ways appropriate to the user, context, and interaction purpose. This includes interaction properties such as framing, tone, perceived authority, responsiveness, uncertainty handling, and conversational guidance. Existing evaluations capture parts of this problem but rarely assess these psychological effects directly. The authors outline a conceptual framework for psychological competence and its core domains, and describe how it may be assessed through scenario-based probes, structured human evaluation, and model-assisted evaluation methods. They argue that psychological competence should become a core consideration for model providers, deploying organizations, researchers, and regulators concerned with the real-world effects of human-facing AI systems. The paper appears on arXiv.
Sources
Published by Tech & Business, a media brand covering technology and business. This story was sourced from arXiv and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
AI Products
AI Products

SoftBank seeks up to $100 billion from Gulf investors for AI

SoftBank Group Corp. is seeking to raise as much as $100 billion from investors in the Gulf region to help finance an expansion of the company's investments in artificial intelligence, the Financial Times has reported....

AI Science
AI Science

Nullify preprint reports selective LLM forgetting without weight updates

Researchers report in a preprint that Nullify, a method designed to suppress specific memorized information in large language models, matches or surpasses established forgetting baselines on TOFU and MUSE while preserving model ut...

AI Science
AI Science

Preprint reports schema-free generation of valid enterprise test data

Researchers report in an arXiv preprint that their Generalist Populator agent generated enterprise data with 100% constraint satisfaction and 0.88 average marginal fidelity across ten simulated environments without accessing datab...

AI Science
AI Science

Preprint reports task-completion gains from agent-generated interfaces

Researchers report in an arXiv preprint that training an agent to generate interactive interfaces improved a 4B model's Pass@3 task-completion score from 9.33% to 58.00%. Their GenUI-Harness pairs an agent that retrieves informati...