Saved in:
| Main Authors: | Wang, Linna, You, Zhixuan, Zhang, Qihui, Wen, Jiunan, Shi, Ji, Chen, Yimin, Wang, Yusen, Ding, Fanqi, Feng, Ziliang, Lu, Li |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2511.07127 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
REACT: Representation Extraction And Controllable Tuning to Overcome Overfitting in LLM Knowledge Editing
by: Zhong, Haitian, et al.
Published: (2025)
by: Zhong, Haitian, et al.
Published: (2025)
Task-Driven Causal Feature Distillation: Towards Trustworthy Risk Prediction
by: Chu, Zhixuan, et al.
Published: (2023)
by: Chu, Zhixuan, et al.
Published: (2023)
Train Yourself as an LLM: Exploring Effects of AI Literacy on Persuasion via Role-playing LLM Training
by: Fan, Qihui, et al.
Published: (2026)
by: Fan, Qihui, et al.
Published: (2026)
Sentinel REACT
by: Woodworth, Michael
Published: (2025)
by: Woodworth, Michael
Published: (2025)
Sentinel REACT
by: Woodworth, Michael
Published: (2025)
by: Woodworth, Michael
Published: (2025)
CreditARF: A Framework for Corporate Credit Rating with Annual Report and Financial Feature Integration
by: Shi, Yumeng, et al.
Published: (2025)
by: Shi, Yumeng, et al.
Published: (2025)
LLM Sensitivity Evaluation Framework for Clinical Diagnosis
by: Yan, Chenwei, et al.
Published: (2025)
by: Yan, Chenwei, et al.
Published: (2025)
CausalFlip: A Benchmark for LLM Causal Judgment Beyond Semantic Matching
by: Wang, Yuzhe, et al.
Published: (2026)
by: Wang, Yuzhe, et al.
Published: (2026)
Cocktail: A Comprehensive Information Retrieval Benchmark with LLM-Generated Documents Integration
by: Dai, Sunhao, et al.
Published: (2024)
by: Dai, Sunhao, et al.
Published: (2024)
LLM-Extracted Covariates for Clinical Causal Inference: Rethinking Integration Strategies
by: Liu, Lei, et al.
Published: (2026)
by: Liu, Lei, et al.
Published: (2026)
OracleProto: A Reproducible Framework for Benchmarking LLM Native Forecasting via Knowledge Cutoff and Temporal Masking
by: Ma, Yiding, et al.
Published: (2026)
by: Ma, Yiding, et al.
Published: (2026)
Benchmark^2: Systematic Evaluation of LLM Benchmarks
by: Qian, Qi, et al.
Published: (2026)
by: Qian, Qi, et al.
Published: (2026)
Intelligent Depression Prevention via LLM-Based Dialogue Analysis: Overcoming the Limitations of Scale-Dependent Diagnosis through Precise Emotional Pattern Recognition
by: Zhong, Zhenguang, et al.
Published: (2025)
by: Zhong, Zhenguang, et al.
Published: (2025)
AcademicEval: Live Long-Context LLM Benchmark
by: Zhang, Haozhen, et al.
Published: (2025)
by: Zhang, Haozhen, et al.
Published: (2025)
Understanding LLM Behaviors via Compression: Data Generation, Knowledge Acquisition and Scaling Laws
by: Pan, Zhixuan, et al.
Published: (2025)
by: Pan, Zhixuan, et al.
Published: (2025)
ToMPO: Training LLM Strategic Decision Making from a Multi-Agent Perspective
by: Zhang, Yiwen, et al.
Published: (2025)
by: Zhang, Yiwen, et al.
Published: (2025)
NextQuill: Causal Preference Modeling for Enhancing LLM Personalization
by: Zhao, Xiaoyan, et al.
Published: (2025)
by: Zhao, Xiaoyan, et al.
Published: (2025)
AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation
by: Wang, Junyang, et al.
Published: (2023)
by: Wang, Junyang, et al.
Published: (2023)
Causal Interventional Prediction System for Robust and Explainable Effect Forecasting
by: Chu, Zhixuan, et al.
Published: (2024)
by: Chu, Zhixuan, et al.
Published: (2024)
LLM-GLOBE: A Benchmark Evaluating the Cultural Values Embedded in LLM Output
by: Karinshak, Elise, et al.
Published: (2024)
by: Karinshak, Elise, et al.
Published: (2024)
A Framework for Benchmarking and Aligning Task-Planning Safety in LLM-Based Embodied Agents
by: Huang, Yuting, et al.
Published: (2025)
by: Huang, Yuting, et al.
Published: (2025)
Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration
by: Feng, Shangbin, et al.
Published: (2024)
by: Feng, Shangbin, et al.
Published: (2024)
What Matters in LLM-Based Feature Extractor for Recommender? A Systematic Analysis of Prompts, Models, and Adaptation
by: Shi, Kainan, et al.
Published: (2025)
by: Shi, Kainan, et al.
Published: (2025)
Raccoon: Prompt Extraction Benchmark of LLM-Integrated Applications
by: Wang, Junlin, et al.
Published: (2024)
by: Wang, Junlin, et al.
Published: (2024)
A Course Shared Task on Evaluating LLM Output for Clinical Questions
by: Hou, Yufang, et al.
Published: (2024)
by: Hou, Yufang, et al.
Published: (2024)
Digging Into the Internal: Causality-Based Analysis of LLM Function Calling
by: Ji, Zhenlan, et al.
Published: (2025)
by: Ji, Zhenlan, et al.
Published: (2025)
Enhancing LLM-Based Social Bot via an Adversarial Learning Framework
by: Kong, Fanqi, et al.
Published: (2025)
by: Kong, Fanqi, et al.
Published: (2025)
Glycoprotein N Genotyping in Congenital Cytomegalovirus Infection: Mechanistic Evidence and Clinical Prognostication
by: DuJiang Yang, et al.
Published: (2025)
by: DuJiang Yang, et al.
Published: (2025)
Prognostic Model and Clinical Features for Overall Survival in Pediatric Liposarcoma: A Population‐Based Study
by: Yang Wu, et al.
Published: (2025)
by: Yang Wu, et al.
Published: (2025)
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
by: Long, Xiang, et al.
Published: (2026)
by: Long, Xiang, et al.
Published: (2026)
AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning
by: Ding, Liang
Published: (2026)
by: Ding, Liang
Published: (2026)
How Social is It? A Benchmark for LLMs' Capabilities in Multi-user Multi-turn Social Agent Tasks
by: Wu, Yusen, et al.
Published: (2025)
by: Wu, Yusen, et al.
Published: (2025)
Fed-REACT: Federated Representation Learning for Heterogeneous and Evolving Data
by: Chen, Yiyue, et al.
Published: (2025)
by: Chen, Yiyue, et al.
Published: (2025)
The Evaluation Game: Beyond Static LLM Benchmarking
by: Wang, Paul, et al.
Published: (2026)
by: Wang, Paul, et al.
Published: (2026)
Dual-Valued Functions of Dual Matrices with Applications in Causal Emergence
by: Wei, Tong, et al.
Published: (2024)
by: Wei, Tong, et al.
Published: (2024)
REACT: Multi Robot Energy-Aware Orchestrator for Indoor Search and Rescue Critical Tasks
by: Maresca, Fabio, et al.
Published: (2025)
by: Maresca, Fabio, et al.
Published: (2025)
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
by: Chen, Dongping, et al.
Published: (2024)
by: Chen, Dongping, et al.
Published: (2024)
PHM-Bench: A Domain-Specific Benchmarking Framework for Systematic Evaluation of Large Models in Prognostics and Health Management
by: Yang, Puyu, et al.
Published: (2025)
by: Yang, Puyu, et al.
Published: (2025)
AgentRx: A Benchmark Study of LLM Agents for Multimodal Clinical Prediction Tasks
by: Jorf, Baraa Al, et al.
Published: (2026)
by: Jorf, Baraa Al, et al.
Published: (2026)
A General Benchmark Framework is Dynamic Graph Neural Network Need
by: Zhang, Yusen
Published: (2024)
by: Zhang, Yusen
Published: (2024)
Similar Items
-
REACT: Representation Extraction And Controllable Tuning to Overcome Overfitting in LLM Knowledge Editing
by: Zhong, Haitian, et al.
Published: (2025) -
Task-Driven Causal Feature Distillation: Towards Trustworthy Risk Prediction
by: Chu, Zhixuan, et al.
Published: (2023) -
Train Yourself as an LLM: Exploring Effects of AI Literacy on Persuasion via Role-playing LLM Training
by: Fan, Qihui, et al.
Published: (2026) -
Sentinel REACT
by: Woodworth, Michael
Published: (2025) -
Sentinel REACT
by: Woodworth, Michael
Published: (2025)