NORA-1.5: A Vision-Language-Action Model Trained using World Model- and Action-based Preference Rewards
Fuente:
arXiv
Saved in:
| Main Authors: | Hung, Chia-Yu, Majumder, Navonil, Deng, Haoyuan, Renhang, Liu, Ang, Yankang, Zadeh, Amir, Li, Chuan, Herremans, Dorien, Wang, Ziwei, Poria, Soujanya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks
by: Hung, Chia-Yu, et al.
Published: (2025)
by: Hung, Chia-Yu, et al.
Published: (2025)
JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment
by: Liu, Renhang, et al.
Published: (2025)
by: Liu, Renhang, et al.
Published: (2025)
Inference Time Alignment with Reward-Guided Tree Search
by: Hung, Chia-Yu, et al.
Published: (2024)
by: Hung, Chia-Yu, et al.
Published: (2024)
10 Open Challenges Steering the Future of Vision-Language-Action Models
by: Poria, Soujanya, et al.
Published: (2025)
by: Poria, Soujanya, et al.
Published: (2025)
Mustango: Toward Controllable Text-to-Music Generation
by: Melechovsky, Jan, et al.
Published: (2023)
by: Melechovsky, Jan, et al.
Published: (2023)
Video2Music: Suitable Music Generation from Videos using an Affective Multimodal Transformer model
by: Kang, Jaeyong, et al.
Published: (2023)
by: Kang, Jaeyong, et al.
Published: (2023)
Lessons from Training Grounded LLMs with Verifiable Rewards
by: Sim, Shang Hong, et al.
Published: (2025)
by: Sim, Shang Hong, et al.
Published: (2025)
Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization
by: Majumder, Navonil, et al.
Published: (2024)
by: Majumder, Navonil, et al.
Published: (2024)
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision
by: Pala, Tej Deep, et al.
Published: (2025)
by: Pala, Tej Deep, et al.
Published: (2025)
Leveraging LLM Embeddings for Cross Dataset Label Alignment and Zero Shot Music Emotion Prediction
by: Liu, Renhang, et al.
Published: (2024)
by: Liu, Renhang, et al.
Published: (2024)
TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
by: Hung, Chia-Yu, et al.
Published: (2024)
by: Hung, Chia-Yu, et al.
Published: (2024)
LLMs Can't Handle Peer Pressure: Crumbling under Multi-Agent Social Interactions
by: Song, Maojia, et al.
Published: (2025)
by: Song, Maojia, et al.
Published: (2025)
Demystifying deep search: a holistic evaluation with hint-free multi-hop questions and factorised metrics
by: Song, Maojia, et al.
Published: (2025)
by: Song, Maojia, et al.
Published: (2025)
JamendoMaxCaps: A Large Scale Music-caption Dataset with Imputed Metadata
by: Roy, Abhinaba, et al.
Published: (2025)
by: Roy, Abhinaba, et al.
Published: (2025)
Evaluating LLMs' Mathematical and Coding Competency through Ontology-guided Interventions
by: Hong, Pengfei, et al.
Published: (2024)
by: Hong, Pengfei, et al.
Published: (2024)
Aligning Generative Music AI with Human Preferences: Methods and Challenges
by: Herremans, Dorien, et al.
Published: (2025)
by: Herremans, Dorien, et al.
Published: (2025)
PromptDistill: Query-based Selective Token Retention in Intermediate Layers for Efficient Large Language Model Inference
by: Jin, Weisheng, et al.
Published: (2025)
by: Jin, Weisheng, et al.
Published: (2025)
Improving Text-To-Audio Models with Synthetic Captions
by: Kong, Zhifeng, et al.
Published: (2024)
by: Kong, Zhifeng, et al.
Published: (2024)
Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse
by: Song, Maojia, et al.
Published: (2024)
by: Song, Maojia, et al.
Published: (2024)
Towards Unified Music Emotion Recognition across Dimensional and Categorical Models
by: Kang, Jaeyong, et al.
Published: (2025)
by: Kang, Jaeyong, et al.
Published: (2025)
Are We There Yet? A Brief Survey of Music Emotion Prediction Datasets, Models and Outstanding Challenges
by: Kang, Jaeyong, et al.
Published: (2024)
by: Kang, Jaeyong, et al.
Published: (2024)
OffTopicEval: When Large Language Models Enter the Wrong Chat, Almost Always!
by: Lei, Jingdi, et al.
Published: (2025)
by: Lei, Jingdi, et al.
Published: (2025)
Training Vision-Language Process Reward Models for Test-Time Scaling in Multimodal Reasoning: Key Insights and Lessons Learned
by: Ong, Brandon, et al.
Published: (2025)
by: Ong, Brandon, et al.
Published: (2025)
DisfluencySpeech -- Single-Speaker Conversational Speech Dataset with Paralanguage
by: Wang, Kyra, et al.
Published: (2024)
by: Wang, Kyra, et al.
Published: (2024)
PreBit -- A multimodal model with Twitter FinBERT embeddings for extreme price movement prediction of Bitcoin
by: Zou, Yanzhao, et al.
Published: (2022)
by: Zou, Yanzhao, et al.
Published: (2022)
DeepUnifiedMom: Unified Time-series Momentum Portfolio Construction via Multi-Task Learning with Multi-Gate Mixture of Experts
by: Ong, Joel, et al.
Published: (2024)
by: Ong, Joel, et al.
Published: (2024)
IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model
by: Jiang, Anqing, et al.
Published: (2025)
by: Jiang, Anqing, et al.
Published: (2025)
Can-Do! A Dataset and Neuro-Symbolic Grounded Framework for Embodied Planning with Large Multimodal Models
by: Chia, Yew Ken, et al.
Published: (2024)
by: Chia, Yew Ken, et al.
Published: (2024)
Towards Robust Instruction Tuning on Multimodal Large Language Models
by: Han, Wei, et al.
Published: (2024)
by: Han, Wei, et al.
Published: (2024)
APEX: Large-scale Multi-task Aesthetic-Informed Popularity Prediction for AI-Generated Music
by: Husain, Jaavid Aktar, et al.
Published: (2026)
by: Husain, Jaavid Aktar, et al.
Published: (2026)
Forecasting Bitcoin volatility spikes from whale transactions and CryptoQuant data using Synthesizer Transformer models
by: Herremans, Dorien, et al.
Published: (2022)
by: Herremans, Dorien, et al.
Published: (2022)
Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic
by: Bhardwaj, Rishabh, et al.
Published: (2024)
by: Bhardwaj, Rishabh, et al.
Published: (2024)
Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training
by: Melechovsky, Jan, et al.
Published: (2024)
by: Melechovsky, Jan, et al.
Published: (2024)
From Perception to Action: An Interactive Benchmark for Vision Reasoning
by: Wu, Yuhao, et al.
Published: (2026)
by: Wu, Yuhao, et al.
Published: (2026)
Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning
by: Sun, Qi, et al.
Published: (2024)
by: Sun, Qi, et al.
Published: (2024)
Are Language Models Puzzle Prodigies? Algorithmic Puzzles Unveil Serious Challenges in Multimodal Reasoning
by: Ghosal, Deepanway, et al.
Published: (2024)
by: Ghosal, Deepanway, et al.
Published: (2024)
Prevailing Research Areas for Music AI in the Era of Foundation Models
by: Wei, Megan, et al.
Published: (2024)
by: Wei, Megan, et al.
Published: (2024)
The Jumping Reasoning Curve? Tracking the Evolution of Reasoning Performance in GPT-[n] and o-[n] Models on Multimodal Puzzles
by: Toh, Vernon Y. H., et al.
Published: (2025)
by: Toh, Vernon Y. H., et al.
Published: (2025)
DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling
by: Deep, Pala Tej, et al.
Published: (2024)
by: Deep, Pala Tej, et al.
Published: (2024)
VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search
by: Guo, Wenkai, et al.
Published: (2025)
by: Guo, Wenkai, et al.
Published: (2025)
Similar Items
-
NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks
by: Hung, Chia-Yu, et al.
Published: (2025) -
JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment
by: Liu, Renhang, et al.
Published: (2025) -
Inference Time Alignment with Reward-Guided Tree Search
by: Hung, Chia-Yu, et al.
Published: (2024) -
10 Open Challenges Steering the Future of Vision-Language-Action Models
by: Poria, Soujanya, et al.
Published: (2025) -
Mustango: Toward Controllable Text-to-Music Generation
by: Melechovsky, Jan, et al.
Published: (2023)