A2ASecBench: A Protocol-Aware Security Benchmark for Agent-to-Agent Multi-Agent Systems
Publications
2026
Are All Prompt Components Value-Neutral? Understanding the Heterogeneous Adversarial Robustness of Dissected Prompt in Large Language Models
Emerging from Ground: Addressing Intent Deviation in Tool-Using Agents via Deriving Real Calls into Virtual Trajectories
Beyond History Mirroring: Profile-Driven Bidirectional Intent Evolution for Personalized Query Recommendation
Beyond Accuracy: Comprehensive Alignment for AI-Driven Molecular Design
A Versatile Method for Accurately Predicting Electronic Absorption Spectra of Tetrapyrrole Macrocycles
2025
AI Benchmark Democratization and Carpentry
SciSafeEval: A Comprehensive Benchmark for Safety Alignment of Large Language Models in Scientific Tasks
Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems
SFSWTS: A Spatial-Frequency Shifted Windows and Time Self-Attention Network for EEG Emotion Recognition
A Survey of Vibe Coding with Large Language Models
From Droplets to Diagnosis: AI-Driven Imaging and System Integration in Digital Nucleic Acid Amplification Testing
Technical Reports
A Robust, Defensible, and Reproducible Methodology for Benchmarking Single-Turn Jailbreak Attacks on Large Language Models
Agentic Product Maturity Ladder V0.1
AILuminate Security: Introducing v0.5 of the Jailbreak Benchmark from MLCommons
AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons