2025年人工智能安全指数报告_114页_1mb
报告摘要
AI Safety Index Summary
Core Content
The AI Safety Index by the Future of Life Institute (FLI) is an independent evaluation of eight leading AI companies, assessing their efforts in managing immediate harms and catastrophic risks from advanced AI systems. The third iteration of the index, released in Winter 2025, highlights the growing capabilities of AI systems and the urgent need for robust safety practices and transparent risk management. The index evaluates companies across six critical domains: Risk Assessment, Current Harms, Safety Frameworks, Existential Safety, Governance & Accountability, and Information Sharing.
Main Findings
- Top Performers: Anthropic, OpenAI, and Google DeepMind maintain their top positions, with Anthropic leading in all domains. However, they still face gaps in concrete safeguards, independent oversight, and credible long-term risk-management strategies.
- Performance Gaps: xAI, Z.ai, Meta, DeepSeek, and Alibaba Cloud score lower, with existential safety being a major structural failure across the industry. No company scored above a D in this domain for the second consecutive year.
- Improvement Signs: Some companies, like xAI and Z.ai, have made progress in publishing safety frameworks and increasing transparency. However, these frameworks remain limited in scope, measurability, and independent oversight.
- Regulatory Influence: Chinese companies are noted for stronger baseline accountability due to domestic regulations such as content labeling and incident reporting. However, they still lack international transparency and voluntary commitments.
- Standards Gap: Companies are lagging behind emerging standards like the EU AI Code of Practice. They fail to meet basic requirements such as independent oversight, transparent threat modeling, measurable thresholds, and clear mitigation triggers.
Key Domains and Scores
| Domain | Anthropic | OpenAI | Google DeepMind | xAI | Z.ai | Meta | DeepSeek | Alibaba Cloud |
|---|---|---|---|---|---|---|---|---|
| Overall Grade | C+ | C+ | C | D | D | D | D | D- |
| Score | 2.67 | 2.31 | 2.08 | 1.17 | 1.12 | 1.10 | 1.02 | 0.98 |
| Risk Assessment | B | B | C+ | D | D+ | D | D | D |
| Current Harms | C+ | C- | C | F | D | D+ | D+ | D+ |
| Safety Frameworks | C+ | C+ | C+ | D+ | D- | D+ | F | F |
| Existential Safety | D | D | D | F | F | F | F | F |
| Governance & Accountability | B- | C+ | C- | D | D | D | D | D+ |
| Information Sharing | A- | B | C | C | C- | D- | C- | D+ |
Company Progress Highlights and Improvement Recommendations
Anthropic
- Highlights: Increased transparency, improved governance, and support for state and international AI safety initiatives.
- Recommendations: Make thresholds and safeguards more concrete and measurable. Strengthen evaluation methodology and independence.
OpenAI
- Highlights: Documented a broader risk assessment process.
- Recommendations: Make safety framework thresholds measurable and enforceable. Increase transparency and external oversight. Improve model robustness and reduce adversarial behavior.
Google DeepMind
- Highlights: Improved transparency and governance mechanisms.
- Recommendations: Strengthen risk-assessment rigor and independence. Define measurable criteria and clarify decision-making authority.
xAI
- Highlights: Formalized and published its frontier AI safety framework.
- Recommendations: Improve breadth, rigor, and independence of risk assessments. Clarify risk management framework and allow more pre-deployment testing.
Z.ai
- Highlights: Took a step toward external oversight by allowing third-party evaluators to publish results and deferring to external authorities for emergency response.
- Recommendations: Publicize full safety framework and governance structure. Improve model robustness and trustworthiness. Establish a whistleblower policy.
Meta
- Highlights: Published a frontier AI safety framework with clear thresholds and risk modeling.
- Recommendations: Improve risk assessment and safety evaluations. Strengthen internal safety governance. Foster a culture of serious risk consideration. Improve information sharing.
DeepSeek
- Highlights: Employees have become more vocal about AI risks, and the company has contributed to standard-setting.
- Recommendations: Establish and publish a foundational safety framework. Improve model robustness and trustworthiness. Establish a whistleblower policy and bug bounty program.
Alibaba Cloud
- Highlights: Contributed to national watermarking standards.
- Recommendations: Establish and publish a foundational safety framework. Improve model robustness and trustworthiness. Establish a whistleblower policy. Improve information sharing.
Methodology
The AI Safety Index evaluates companies on 35 indicators across six domains. The methodology includes:
- Indicator Selection: Based on regulatory and voluntary commitments, with a focus on risk assessment, current harms, safety frameworks, existential safety, governance, and information sharing.
- Company Selection: Includes the top 8 AI companies, with Anthropic, OpenAI, Google DeepMind, xAI, Z.ai, Meta, DeepSeek, and Alibaba Cloud.
- Evidence Collection: Combines publicly available materials (model cards, research papers, benchmark results) with targeted company surveys to address transparency gaps.
- Grading: Conducted by an independent panel of AI researchers and governance experts. Final scores are averaged expert assessments.
Independent Review Panel
The index is evaluated by a panel of eight distinguished AI experts, including:
- David Krueger – AI safety and alignment researcher at University of Montreal.
- Dylan Hadfield-Menell – AI safety and accountability researcher at MIT.
- Stuart Russell – Renowned AI researcher and co-author of the standard AI textbook.
- Sharon Li – AI safety and reliability researcher at University of Wisconsin-Madison.
- Jessica Newman – AI security expert at UC Berkeley.
- Sneha Revanur – Founder of Encode, a youth-led AI ethics organization.
- Tegan Maharaj – Responsible AI researcher at HEC Montréal.
- Yi Zeng – AI professor at the Chinese Academy of Sciences and leader in AI safety governance.
Conclusion
The AI Safety Index highlights that while some companies are making progress, the frontier AI ecosystem still lacks credible control plans and robust safety practices. The gap between capability and safety continues to widen, raising concerns about the long-term controllability and alignment of increasingly powerful AI systems. The index serves as a tool for transparency, comparative analysis, and identifying critical gaps in AI safety practices. It emphasizes the need for concrete safeguards, measurable thresholds, and independent oversight to ensure AI systems are aligned with human values and safe to deploy.
试读结束,高清完整版pdf/doc/ppt,请点下载