Back to feed
News Story
机器之心
1 sources

ECNU Team Releases Finance-Native Foundation Model Mint-Agent, Outperforming GPT-5.6 and Claude Opus 4.8

The team at East China Normal University's Institute of Intelligent Finance has released Mint-Agent, a finance-native agentic foundation model family including the 9B Mint-Cu and 27B Mint-Ag. On the FinanceAgentBench v2 benchmark, Mint-Ag scored 60.49%, surpassing GPT-5.6-Sol and Claude Opus 4.8, while its per-question inference cost is only about 22% and 10% of the two, respectively. Commercial orders have already exceeded 100 million yuan, marking a shift from capability validation to large-scale financial scenario validation.

SynthePulse Insight · AI deep readingMembers

Mint-Agent: How a Finance-Native Foundation Model Turns 'Auditability' into Infrastructure

Version 1 · 1 source

The East China Normal University Institute of AI Finance releases Mint-Agent, a finance-native Agentic Foundation Model family. The 27B Mint-Ag scores 60.49% on FinanceAgentBench v2, surpassing GPT-5.6-Sol and Claude Opus 4.8, with per-question inference costs of only about 22% and 10% of those models, respectively. Commercial orders have exceeded 100 million RMB. Its core is not stronger Q&A, but integrating source, time, numbers, and operations into a single recoverable chain, making every conclusion verifiable.

  • The Mint-Agent family includes the 9B Mint-Cu and the 27B Mint-Ag, positioned as finance-native Agentic Foundation Models.
  • On FinanceAgentBench v2, Mint-Ag scores 60.49%, surpassing GPT-5.6-Sol and Claude Opus 4.8, with per-question inference costs of about 22% and 10% of those models, respectively.
  • The model integrates source, time, numbers, and operations into a single recoverable chain, allowing reliability to be traced to specific steps in evidence, extraction, or calculation.
Open section navigationPerformance and Cost: Not Just a Stronger Model

Performance and Cost: Not Just a Stronger Model

The key data from Mint-Agent's release: on FinanceAgentBench v2, the 27B Mint-Ag scores 60.49%, surpassing GPT-5.6-Sol and Claude Opus 4.8, while per-question inference costs are only about 22% and 10% of those models, respectively. These numbers directly address a critical question in financial AI: does performance improvement necessarily come at the cost of higher execution costs? Mint-Agent's answer is that by coordinating financial reasoning, long-horizon execution, and evidence management, higher accuracy can be achieved without increasing the inference budget.

The 9B Mint-Cu demonstrates the efficiency potential of compact models in financial search and execution. The team emphasizes that the real hurdle is not 'knowing' but 'completing'—real financial research requires the model to find authoritative materials, identify the correct reporting period, extract metrics, unify units, perform calculations, cross-verify across sources, and form conclusions. If any step deviates, the final text may 'look reasonable but actually be wrong.'

Free for now

Read the full analysis

4 more sections of analysis, plus the full takeaway

Loading

Credibility boundary

This article's information primarily comes from a report by Jiqizhixin on the release of Mint-Agent, which is a second-hand source. All performance data, cost comparisons, and commercial order amounts are as claimed by the source and have not been independently verified. Technical details (such as MintHarness, Evidence Ledger, and training paths) come from the team's published technical reports and project website, also as source claims. Readers should treat them as vendor claims rather than confirmed facts.

Primary report

机器之心

Primary source