0 citations

Evaluating Frontier Models for Dangerous Capabilities

arXiv (Cornell University)2024

Citations Over Time

Abstract

To understand the risks posed by a new AI system, we must understand what it can and cannot do. Building on prior work, we introduce a programme of new "dangerous capability" evaluations and pilot them on Gemini 1.0 models. Our evaluations cover four areas: (1) persuasion and deception; (2) cyber-security; (3) self-proliferation; and (4) self-reasoning. We do not find evidence of strong dangerous capabilities in the models we evaluated, but we flag early warning signs. Our goal is to help advance a rigorous science of dangerous capability evaluation, in preparation for future models.

Related Papers

→ Frontier Cities: Encounters at the Crossroads of Empire(2014)8 cited
→ The Wild, Wild Web: The Mythic American West and the Electronic Frontier(2000)24 cited
→ Южный и Восточный Фронтир Россиив XVI – XVIII Веках(2003)10 cited
Research on Problems concerning Frontier Control of Southwest Territory(2008)
Susquehanna Chorale Spring Concert "Roots and Wings"(2017)