Hummer: Towards Limited Competitive Preference Dataset
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiang, Li, Wu, Yusen, Xiong, Junwu, Ruan, Jingqing, Ding, Yichuan, Guo, Qingpei, Wen, Zujie, Zhou, Jun, Deng, Xiaotie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Social is It? A Benchmark for LLMs' Capabilities in Multi-user Multi-turn Social Agent Tasks
von: Wu, Yusen, et al.
Veröffentlicht: (2025)
von: Wu, Yusen, et al.
Veröffentlicht: (2025)
MALLES: A Multi-agent LLMs-based Economic Sandbox with Consumer Preference Alignment
von: Wu, Yusen, et al.
Veröffentlicht: (2026)
von: Wu, Yusen, et al.
Veröffentlicht: (2026)
DeepRule: An Integrated Framework for Automated Business Rule Generation via Deep Predictive Modeling and Hybrid Search Optimization
von: Wu, Yusen, et al.
Veröffentlicht: (2025)
von: Wu, Yusen, et al.
Veröffentlicht: (2025)
Implementing Long Text Style Transfer with LLMs through Dual-Layered Sentence and Paragraph Structure Extraction and Mapping
von: Wu, Yusen, et al.
Veröffentlicht: (2025)
von: Wu, Yusen, et al.
Veröffentlicht: (2025)
HCAG: Hierarchical Abstraction and Retrieval-Augmented Generation on Theoretical Repositories with LLMs
von: Wu, Yusen, et al.
Veröffentlicht: (2026)
von: Wu, Yusen, et al.
Veröffentlicht: (2026)
How Large Language Models Need Symbolism
von: Deng, Xiaotie, et al.
Veröffentlicht: (2025)
von: Deng, Xiaotie, et al.
Veröffentlicht: (2025)
Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers
von: Sun, Haoran, et al.
Veröffentlicht: (2025)
von: Sun, Haoran, et al.
Veröffentlicht: (2025)
OrdMoE: Preference Alignment via Hierarchical Expert Group Ranking in Multimodal Mixture-of-Experts LLMs
von: Gao, Yuting, et al.
Veröffentlicht: (2025)
von: Gao, Yuting, et al.
Veröffentlicht: (2025)
Will AI Trade? A Computational Inversion of the No-Trade Theorem
von: Li, Hanyu, et al.
Veröffentlicht: (2025)
von: Li, Hanyu, et al.
Veröffentlicht: (2025)
X-Light: Cross-City Traffic Signal Control Using Transformer on Transformer as Meta Multi-Agent Reinforcement Learner
von: Jiang, Haoyuan, et al.
Veröffentlicht: (2024)
von: Jiang, Haoyuan, et al.
Veröffentlicht: (2024)
RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training
von: Sun, Haoran, et al.
Veröffentlicht: (2026)
von: Sun, Haoran, et al.
Veröffentlicht: (2026)
AI Agents Under Threat: A Survey of Key Security Challenges and Future Pathways
von: Deng, Zehang, et al.
Veröffentlicht: (2024)
von: Deng, Zehang, et al.
Veröffentlicht: (2024)
An Information-Theoretic Criterion for Efficient Data Synthesis
von: Li, Hanyu, et al.
Veröffentlicht: (2026)
von: Li, Hanyu, et al.
Veröffentlicht: (2026)
PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling
von: Jian, Ai, et al.
Veröffentlicht: (2025)
von: Jian, Ai, et al.
Veröffentlicht: (2025)
Learning Top-k Subtask Planning Tree based on Discriminative Representation Pre-training for Decision Making
von: Ruan, Jingqing, et al.
Veröffentlicht: (2023)
von: Ruan, Jingqing, et al.
Veröffentlicht: (2023)
CoSLight: Co-optimizing Collaborator Selection and Decision-making to Enhance Traffic Signal Control
von: Ruan, Jingqing, et al.
Veröffentlicht: (2024)
von: Ruan, Jingqing, et al.
Veröffentlicht: (2024)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
von: Deng, Yichuan, et al.
Veröffentlicht: (2024)
von: Deng, Yichuan, et al.
Veröffentlicht: (2024)
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models
von: Liu, Qi, et al.
Veröffentlicht: (2025)
von: Liu, Qi, et al.
Veröffentlicht: (2025)
Towards Data-Centric RLHF: Simple Metrics for Preference Dataset Comparison
von: Shen, Judy Hanwen, et al.
Veröffentlicht: (2024)
von: Shen, Judy Hanwen, et al.
Veröffentlicht: (2024)
M2-omni: Advancing Omni-MLLM for Comprehensive Modality Support with Competitive Performance
von: Guo, Qingpei, et al.
Veröffentlicht: (2025)
von: Guo, Qingpei, et al.
Veröffentlicht: (2025)
AdaptiveLoad: Towards Efficient Video Diffusion Transformer Training
von: Guo, Yucheng, et al.
Veröffentlicht: (2026)
von: Guo, Yucheng, et al.
Veröffentlicht: (2026)
Efficient Reinforcement Learning for Large Language Models with Intrinsic Exploration
von: Sun, Yan, et al.
Veröffentlicht: (2025)
von: Sun, Yan, et al.
Veröffentlicht: (2025)
InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning
von: Yan, Yuchen, et al.
Veröffentlicht: (2026)
von: Yan, Yuchen, et al.
Veröffentlicht: (2026)
USM: Unbiased Survey Modeling for Limiting Negative User Experiences in Recommendation Systems
von: Yu, Chenghui, et al.
Veröffentlicht: (2024)
von: Yu, Chenghui, et al.
Veröffentlicht: (2024)
Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems
von: Cui, Tianyu, et al.
Veröffentlicht: (2024)
von: Cui, Tianyu, et al.
Veröffentlicht: (2024)
CompetEvo: Towards Morphological Evolution from Competition
von: Huang, Kangyao, et al.
Veröffentlicht: (2024)
von: Huang, Kangyao, et al.
Veröffentlicht: (2024)
Enhancing Stochastic Gradient Descent: A Unified Framework and Novel Acceleration Methods for Faster Convergence
von: Deng, Yichuan, et al.
Veröffentlicht: (2024)
von: Deng, Yichuan, et al.
Veröffentlicht: (2024)
CLLMRec: LLM-powered Cognitive-Aware Concept Recommendation via Semantic Alignment and Prerequisite Knowledge Distillation
von: Xiong, Xiangrui, et al.
Veröffentlicht: (2025)
von: Xiong, Xiangrui, et al.
Veröffentlicht: (2025)
FlexPose: Pose Distribution Adaptation with Limited Guidance
von: Wang, Zixiao, et al.
Veröffentlicht: (2024)
von: Wang, Zixiao, et al.
Veröffentlicht: (2024)
Towards Acyclic Preference Evaluation of Language Models via Multiple Evaluators
von: Hu, Zhengyu, et al.
Veröffentlicht: (2024)
von: Hu, Zhengyu, et al.
Veröffentlicht: (2024)
Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs
von: Xuan, Shiyu, et al.
Veröffentlicht: (2023)
von: Xuan, Shiyu, et al.
Veröffentlicht: (2023)
A General Benchmark Framework is Dynamic Graph Neural Network Need
von: Zhang, Yusen
Veröffentlicht: (2024)
von: Zhang, Yusen
Veröffentlicht: (2024)
Sword: Style-Robust World Models as Simulators via Dynamic Latent Bootstrapping for VLA Policy Post-Training
von: Gao, Jiaxuan, et al.
Veröffentlicht: (2026)
von: Gao, Jiaxuan, et al.
Veröffentlicht: (2026)
Improving VTE Identification through Language Models from Radiology Reports: A Comparative Study of Mamba, Phi-3 Mini, and BERT
von: Deng, Jamie, et al.
Veröffentlicht: (2024)
von: Deng, Jamie, et al.
Veröffentlicht: (2024)
Discovering Expert-Level Nash Equilibrium Algorithms with Large Language Models
von: Li, Hanyu, et al.
Veröffentlicht: (2025)
von: Li, Hanyu, et al.
Veröffentlicht: (2025)
Empowering Large Language Models for Textual Data Augmentation
von: Li, Yichuan, et al.
Veröffentlicht: (2024)
von: Li, Yichuan, et al.
Veröffentlicht: (2024)
IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis
von: Li, Hanyu, et al.
Veröffentlicht: (2025)
von: Li, Hanyu, et al.
Veröffentlicht: (2025)
CPRet: A Dataset, Benchmark, and Model for Retrieval in Competitive Programming
von: Deng, Han, et al.
Veröffentlicht: (2025)
von: Deng, Han, et al.
Veröffentlicht: (2025)
Memory-Augmented LLM-based Multi-Agent System for Automated Feature Generation on Tabular Data
von: Dong, Fengxian, et al.
Veröffentlicht: (2026)
von: Dong, Fengxian, et al.
Veröffentlicht: (2026)
TRUST-SQL: Tool-Integrated Multi-Turn Reinforcement Learning for Text-to-SQL over Unknown Schemas
von: Jian, Ai, et al.
Veröffentlicht: (2026)
von: Jian, Ai, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
How Social is It? A Benchmark for LLMs' Capabilities in Multi-user Multi-turn Social Agent Tasks
von: Wu, Yusen, et al.
Veröffentlicht: (2025) -
MALLES: A Multi-agent LLMs-based Economic Sandbox with Consumer Preference Alignment
von: Wu, Yusen, et al.
Veröffentlicht: (2026) -
DeepRule: An Integrated Framework for Automated Business Rule Generation via Deep Predictive Modeling and Hybrid Search Optimization
von: Wu, Yusen, et al.
Veröffentlicht: (2025) -
Implementing Long Text Style Transfer with LLMs through Dual-Layered Sentence and Paragraph Structure Extraction and Mapping
von: Wu, Yusen, et al.
Veröffentlicht: (2025) -
HCAG: Hierarchical Abstraction and Retrieval-Augmented Generation on Theoretical Repositories with LLMs
von: Wu, Yusen, et al.
Veröffentlicht: (2026)