GenAI is evolving at an unprecedented pace, with frequent releases of new large language models (LLMs) featuring performance improvements, efficiency gains, and new capabilities. Developers, researchers, and organizations looking to quickly leverage those model advances face the significant challenge of being able to consistently and reliably evaluate their performance and safety and determine which one is best suited for their use cases. To help address this need, Google DeepMind and Giskard are releasing LMEval, a large model evaluation framework, alongside the Phare Benchmark, an independent multi-lingual security and safety benchmark.
Toward Secure & Trustworthy AI: Independent Benchmarking
| Available Media | |
|---|---|
| Conference | InCyber (InCyber Forum) - 2025 |
| Author | Elie Bursztein |
Recent
ai
Facade: High-Precision Insider Threat Detection Using Deep Contextual Anomaly Detection
publications
Usenix Security 2026
ai
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
publications
NeurIPS 2026
Cybersecurity
DROIDCCT: Cryptographic Compliance Test via Trillion-Scale Measurement
publications
ACSAC 2025