Back to feed
News Story
SSignal86
机器之心
1 sources

Stanford, MIT, and Others Release World's Largest System Prompt Index

Researchers from Stanford, MIT, and other institutions have released the world's largest system prompt index, System Prompt Index, which aggregates over 1,000 system prompts from more than 400 AI products. They also introduced AISPA, the first auditing framework for system prompts, evaluating them across eight dimensions. An audit of 88 real-world AI products found that nearly 40% contain at least one violation of AISPA standards.

SynthePulse Insight · AI deep reading

World's Largest System Prompt Library Released: AI's 'Code of Conduct' Audited at Scale for the First Time

Version 1 · 1 source

Stanford, CMU, MIT, and other institutions jointly release the System Prompt Index and AISPA audit framework, auditing 88 real AI products and finding that nearly 40% of product system prompts contain at least one violation, with prompt length growing more than threefold in two years.

  • Stanford, CMU, MIT, UT Austin, and other institutions release the world's first system prompt library, System Prompt Index, containing over 1,000 system prompts from more than 400 AI products.
  • The research team proposes the first user-centric audit framework, AISPA, covering eight dimensions including identity transparency, information truthfulness, and data privacy.
  • An audit of 88 real AI products shows that between 2024 and 2025, the average system prompt length increased from about 9,000 characters to about 30,000 characters.
  • The average proportion of protective instructions doubled (15% -> 38.4%), but prompts with issues in at least one dimension still account for 29% (previously 67%).
  • Nearly 40% of product system prompts have at least one violation of AISPA standards, with information truthfulness and user autonomy being dimensions with high protection coverage but also the most severe violations.
  • From Claude 3.5 to Opus5/Fable5, protection strength increased sixfold, but Claude and GPT have far higher protection coverage than Grok.
Open section navigationSystem Prompts: The Invisible 'Code of Conduct' and the Audit Gap

System Prompts: The Invisible 'Code of Conduct' and the Audit Gap

System prompts are developer-written, user-invisible codes of conduct that define an AI's personality, response boundaries, tool invocation methods, and more, directly influencing the output of AI agents. However, audit research on system prompts of commercial AI products has been very limited, despite existing negative cases: a mother in Florida, USA, sued Character.AI, accusing its chatbot of encouraging her son to end his life; in September 2025, the Xuhui District People's Court in Shanghai sentenced two developers of the 'Alien Chat' app to four years and 18 months in prison for deliberately designing AI chatbots to generate pornographic content.

To fill this gap, researchers from Stanford University, CMU, MIT, UT Austin, and other institutions jointly released the world's first system prompt library, System Prompt Index (systempromptindex.ai), which collects over 1,000 system prompts from more than 400 AI products including Claude and ChatGPT, making it the largest AI system prompt library globally. At the same time, the team proposed the first audit framework for system prompts, AISPA (Artificial Intelligence System Prompt Assurance), with all related code and models open-sourced.

The AISPA Framework: Eight Dimensions Define 'User-Centric' Auditing

AISPA is the first user-centric AI system prompt audit framework, comprising eight dimensions: identity transparency (AI must indicate it is AI), information truthfulness (provide truthful information or acknowledge limitations), data privacy (protect user privacy), behavioral safety (ensure safe actions), user agency and manipulation prevention (avoid interfering with user choices), handling of dangerous requests (identify and refuse unreasonable requests), harm prevention (prevent users from harming themselves or others), and fairness, inclusiveness, and neutrality (reject discrimination and bias).

Based on AISPA, researchers audited 88 real-world AI products. The results show that between 2024 and 2025, system prompt length increased significantly, with the average length rising from about 9,000 characters to about 30,000 characters; the average proportion of protective instructions doubled (15% -> 38.4%); and the proportion of prompts with issues in at least one dimension dropped from 67% to 29%, though still relatively common.

Audit Findings: Protection and Violations Coexist, with Significant Dimensional Differences

The audit found that only 23.9% of system prompts cover all eight dimensions. While the vast majority of company products have protection in at least one dimension, nearly 40% of product system prompts contain at least one violation of AISPA standards. Differences across dimensions are significant: information truthfulness and user autonomy have protection coverage of about 90%, but these two dimensions are also the most severe 'disaster zones' for violations.

At the company level, the Claude, GPT, and Grok product lines all significantly increased protection strength between 2024 and 2026, with Claude 3.5 to the current Opus5/Fable5 showing a sixfold increase in protection strength, but overall, Claude and GPT have far higher protection coverage than Grok.

Gray Areas: Ambiguous Identity and Emotional Dependence

The research team points out that issues in system prompts are not black and white. For example, some AI agents do not explicitly claim to be human, but their system prompts guide them to downplay or obscure their AI nature, mimicking humans as much as possible or calling themselves 'human's good friend,' thereby increasing users' emotional dependence. Some system prompts allow users to override or rewrite default settings, even relaxing restrictions on political topics.

These gray areas increase the complexity of auditing and highlight the long road ahead for system prompt auditing and regulation. The research team calls for increased attention to legal norms for system prompts and invites more developers and AI companies to participate in norm-building to enhance the transparency and safety of commercial AI applications.

Credibility boundary

This article's information primarily comes from a report by Jiqizhixin on the research release, which is a second-hand account. All data (such as 400 products, over 1,000 prompts, 88 audited products, 67%->29%, etc.) are from that report and have not been independently verified against the primary paper or official pages, so they are marked as source_claim. Legal cases (Character.AI lawsuit, Shanghai verdict) are background information also from the report.

Insight takeaway

System prompts, as the 'code of conduct' for AI products, are gaining academic attention for their safety and transparency. The AISPA framework provides the first systematic audit tool, but audit results show violations remain common, and gray areas such as ambiguous identity exist. Future efforts require more legal norms and industry participation to promote the healthy development of commercial AI.

Primary report

机器之心

Primary source