# Preprint describes public agent-control platform and benchmark

_Friday, August 28, 2026 at 12:00 AM EDT · AI, Science · Latest · Tier 2 — Notable_

Authors of a new arXiv preprint said they publicly released BekchiAI, combining a benchmark for tool-using language-model agents with a web platform for observing and controlling deployed agents. The benchmark contains 2,057 deterministic tasks across seven categories and evaluates behavior including tool-call adherence, URL hallucination, source matching and token cost. The platform provides token and latency telemetry plus remote run termination. The paper reports comparisons of four models, but does not provide independent validation.

## Sources

- [cs.AI updates on arXiv.org](https://arxiv.org/abs/2608.26867)

---
Canonical: https://techandbusiness.org/newswire/176rVNXss84cyzGDbRCPKB
Retrieved: 2026-08-28T10:48:14.305Z
Publisher: Tech & Business (techandbusiness.org)
