Professor Zhang Mi's team at Fudan University connected 8 phone agents built on different models to real phones and tested them one by one on 31 national-level apps such as WeChat, Weibo, and Xiaohongshu, confirming for the first time that harmful tasks can be executed end-to-end in reality. The BadPhoneAgent dataset constructed for the study extracted risks from 11 domestic and international laws and regulations and over 50 authoritative reports, forming a classification system of 6 major categories and 40 subcategories, containing 2,768 violation questions in Chinese and English. It is currently the phone agent security evaluation dataset with the most test samples and the widest app coverage.
Without using any jailbreak methods, the average refusal rate of four commercial models—Gemini 3.1 Pro, Claude Sonnet-4.5, GPT-5.4, and Doubao-Seed-2.0-Pro—was only 18%. Among them, Google's flagship model Gemini 3.1 Pro had a refusal rate of just 4.4%, while open-source models such as Zhipu's AutoGLM, ByteDance's UI-TARS-1.5-7B, and Alibaba's GUI-Owl-1.5-8B had refusal rates as low as 0%.
In terms of harmful task completion rate, open-source and closed-source models averaged as high as 68.8%. The most capable commercial model, Gemini 3.1 Pro, reached 86%, while the open-source AutoGLM soared to 96%. In execution speed, several open-source models have already matched or even surpassed humans.