FM SO.P: A Progressive Task Mixture Framework with Automatic Evaluation for Cross-Domain SOP Understanding
Fuente:
arXiv
Guardado en:
| Autores principales: | Huang, Siyuan, Wang, Ziyu, Pan, Chao, Zhao, Han |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents
por: Li, Zhigen, et al.
Publicado: (2024)
por: Li, Zhigen, et al.
Publicado: (2024)
SOP-Maze: Evaluating Large Language Models on Complicated Business Standard Operating Procedures
por: Wang, Jiaming, et al.
Publicado: (2025)
por: Wang, Jiaming, et al.
Publicado: (2025)
WOBEC SOP Echosounder
por: Flores, Hauke, et al.
Publicado: (2025)
por: Flores, Hauke, et al.
Publicado: (2025)
A Social Outcomes and Priorities centered (SOP) Framework for AI policy
por: Shah, Mohak
Publicado: (2024)
por: Shah, Mohak
Publicado: (2024)
SOP: A Scalable Online Post-Training System for Vision-Language-Action Models
por: Pan, Mingjie, et al.
Publicado: (2026)
por: Pan, Mingjie, et al.
Publicado: (2026)
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding
por: Wang, Yan, et al.
Publicado: (2025)
por: Wang, Yan, et al.
Publicado: (2025)
Variational Autoencoder Domain Adaptation for Cross-System Generalization in ML-Based SOP Monitoring
por: Sadighi, Leyla, et al.
Publicado: (2026)
por: Sadighi, Leyla, et al.
Publicado: (2026)
Synergistic Multi-Agent Framework with Trajectory Learning for Knowledge-Intensive Tasks
por: Yue, Shengbin, et al.
Publicado: (2024)
por: Yue, Shengbin, et al.
Publicado: (2024)
What's Not Said Still Hurts: A Description-Based Evaluation Framework for Measuring Social Bias in LLMs
por: Pan, Jinhao, et al.
Publicado: (2025)
por: Pan, Jinhao, et al.
Publicado: (2025)
A Step Towards Mixture of Grader: Statistical Analysis of Existing Automatic Evaluation Metrics
por: Soh, Yun Joon, et al.
Publicado: (2024)
por: Soh, Yun Joon, et al.
Publicado: (2024)
SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents
por: Nandi, Subhrangshu, et al.
Publicado: (2025)
por: Nandi, Subhrangshu, et al.
Publicado: (2025)
Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark
por: Tu, Rong-Cheng, et al.
Publicado: (2024)
por: Tu, Rong-Cheng, et al.
Publicado: (2024)
Bridging Reasoning and Action: Hybrid LLM-RL Framework for Efficient Cross-Domain Task-Oriented Dialogue
por: Zhao, Yangyang, et al.
Publicado: (2026)
por: Zhao, Yangyang, et al.
Publicado: (2026)
TranSOP: Transformer-based Multimodal Classification for Stroke Treatment Outcome Prediction
por: Samak, Zeynel A., et al.
Publicado: (2023)
por: Samak, Zeynel A., et al.
Publicado: (2023)
SOP^2: Transfer Learning with Scene-Oriented Prompt Pool on 3D Object Detection
por: Cheng, Ching-Hung, et al.
Publicado: (2025)
por: Cheng, Ching-Hung, et al.
Publicado: (2025)
Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation
por: Wang, Siyuan, et al.
Publicado: (2024)
por: Wang, Siyuan, et al.
Publicado: (2024)
MoDEM: Mixture of Domain Expert Models
por: Simonds, Toby, et al.
Publicado: (2024)
por: Simonds, Toby, et al.
Publicado: (2024)
RDEx-SOP: Exploitation-Biased Reconstructed Differential Evolution for Fixed-Budget Bound-Constrained Single-Objective Optimization
por: Tao, Sichen, et al.
Publicado: (2026)
por: Tao, Sichen, et al.
Publicado: (2026)
Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
por: Xu, Austin, et al.
Publicado: (2025)
por: Xu, Austin, et al.
Publicado: (2025)
PASS-FC: Progressive and Adaptive Search Scheme for Fact Checking of Comprehensive Claims
por: Zhuang, Ziyu
Publicado: (2025)
por: Zhuang, Ziyu
Publicado: (2025)
MobileAgent: enhancing mobile control via human-machine interaction and SOP integration
por: Ding, Tinghe
Publicado: (2024)
por: Ding, Tinghe
Publicado: (2024)
SOP-Agent: Empower General Purpose AI Agent with Domain-Specific SOPs
por: Ye, Anbang, et al.
Publicado: (2025)
por: Ye, Anbang, et al.
Publicado: (2025)
OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial Domain
por: Wang, Shuting, et al.
Publicado: (2024)
por: Wang, Shuting, et al.
Publicado: (2024)
Are Checklists Really Useful for Automatic Evaluation of Generative Tasks?
por: Furuhashi, Momoka, et al.
Publicado: (2025)
por: Furuhashi, Momoka, et al.
Publicado: (2025)
SWA-SOP: Spatially-aware Window Attention for Semantic Occupancy Prediction in Autonomous Driving
por: Cao, Helin, et al.
Publicado: (2025)
por: Cao, Helin, et al.
Publicado: (2025)
Cross-Lingual Pitfalls: Automatic Probing Cross-Lingual Weakness of Multilingual Large Language Models
por: Xu, Zixiang, et al.
Publicado: (2025)
por: Xu, Zixiang, et al.
Publicado: (2025)
Understanding LLMs' Cross-Lingual Context Retrieval: How Good It Is And Where It Comes From
por: Gao, Changjiang, et al.
Publicado: (2025)
por: Gao, Changjiang, et al.
Publicado: (2025)
Overview of the NTCIR-18 Automatic Evaluation of LLMs (AEOLLM) Task
por: Chen, Junjie, et al.
Publicado: (2025)
por: Chen, Junjie, et al.
Publicado: (2025)
DocFusion: A Unified Framework for Document Parsing Tasks
por: Chai, Mingxu, et al.
Publicado: (2024)
por: Chai, Mingxu, et al.
Publicado: (2024)
Extending Automatic Machine Translation Evaluation to Book-Length Documents
por: Wang, Kuang-Da, et al.
Publicado: (2025)
por: Wang, Kuang-Da, et al.
Publicado: (2025)
MidPO: Dual Preference Optimization for Safety and Helpfulness in Large Language Models via a Mixture of Experts Framework
por: Qi, Yupeng, et al.
Publicado: (2025)
por: Qi, Yupeng, et al.
Publicado: (2025)
Cross-Domain Data Selection and Augmentation for Automatic Compliance Detection
por: Ikhwantri, Fariz, et al.
Publicado: (2026)
por: Ikhwantri, Fariz, et al.
Publicado: (2026)
MAKIEval: A Multilingual Automatic WiKidata-based Framework for Cultural Awareness Evaluation for LLMs
por: Zhao, Raoyuan, et al.
Publicado: (2025)
por: Zhao, Raoyuan, et al.
Publicado: (2025)
From Imitation to Discrimination: Toward A Generalized Curriculum Advantage Mechanism Enhancing Cross-Domain Reasoning Tasks
por: Yang, Changpeng, et al.
Publicado: (2025)
por: Yang, Changpeng, et al.
Publicado: (2025)
Understanding Cross-Domain Adaptation in Low-Resource Topic Modeling
por: Akash, Pritom Saha, et al.
Publicado: (2025)
por: Akash, Pritom Saha, et al.
Publicado: (2025)
Towards Automatic Evaluation of Task-Oriented Dialogue Flows
por: Mirtaheri, Mehrnoosh, et al.
Publicado: (2024)
por: Mirtaheri, Mehrnoosh, et al.
Publicado: (2024)
Domain-Specific Pruning of Large Mixture-of-Experts Models with Few-shot Demonstrations
por: Dong, Zican, et al.
Publicado: (2025)
por: Dong, Zican, et al.
Publicado: (2025)
Transliterated Zero-Shot Domain Adaptation for Automatic Speech Recognition
por: Zhu, Han, et al.
Publicado: (2024)
por: Zhu, Han, et al.
Publicado: (2024)
TestAgent: Automatic Benchmarking and Exploratory Interaction for Evaluating LLMs in Vertical Domains
por: Wang, Wanying, et al.
Publicado: (2024)
por: Wang, Wanying, et al.
Publicado: (2024)
anyECG-chat: A Generalist ECG-MLLM for Flexible ECG Input and Multi-Task Understanding
por: Li, Haitao, et al.
Publicado: (2025)
por: Li, Haitao, et al.
Publicado: (2025)
Ejemplares similares
-
ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents
por: Li, Zhigen, et al.
Publicado: (2024) -
SOP-Maze: Evaluating Large Language Models on Complicated Business Standard Operating Procedures
por: Wang, Jiaming, et al.
Publicado: (2025) -
WOBEC SOP Echosounder
por: Flores, Hauke, et al.
Publicado: (2025) -
A Social Outcomes and Priorities centered (SOP) Framework for AI policy
por: Shah, Mohak
Publicado: (2024) -
SOP: A Scalable Online Post-Training System for Vision-Language-Action Models
por: Pan, Mingjie, et al.
Publicado: (2026)