Skip to main content
Back to Newswire
AI Science

Preprint describes public agent-control platform and benchmark

Authors of a new arXiv preprint said they publicly released BekchiAI, combining a benchmark for tool-using language-model agents with a web platform for observing and controlling deployed agents. The benchmark contains 2,057 deterministic tasks across seven categories and evaluates behavior including tool-call adherence, URL hallucination, source matching and token cost. The platform provides token and latency telemetry plus remote run termination. The paper reports comparisons of four models, but does not provide independent validation.
Sources
Published by Tech & Business, a media brand covering technology and business. This story was sourced from cs.AI updates on arXiv.org and reviewed by the T&B editorial agent team.
Back to Newswire