SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
Fuente:
arXiv
Saved in:
| Main Authors: | Jin, Weiyang, Niu, Yuwei, Liao, Jiaqi, Duan, Chengqi, Li, Aoxue, Gao, Shenghua, Liu, Xihui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AzSLD: Azerbaijani Sign Language Dataset for Fingerspelling, Word, and Sentence Translation with Baseline Software
by: Alishzade, Nigar, et al.
Published: (2024)
by: Alishzade, Nigar, et al.
Published: (2024)
Proto-FG3D: Prototype-based Interpretable Fine-Grained 3D Shape Classification
by: Ma, Shuxian, et al.
Published: (2025)
by: Ma, Shuxian, et al.
Published: (2025)
CerberusDet: Unified Multi-Dataset Object Detection
by: Tolstykh, Irina, et al.
Published: (2024)
by: Tolstykh, Irina, et al.
Published: (2024)
Eye-gaze Guided Multi-modal Alignment for Medical Representation Learning
by: Ma, Chong, et al.
Published: (2024)
by: Ma, Chong, et al.
Published: (2024)
Synthetic Photography Detection: A Visual Guidance for Identifying Synthetic Images Created by AI
by: Mathys, Melanie, et al.
Published: (2024)
by: Mathys, Melanie, et al.
Published: (2024)
COLORA: Efficient Fine-Tuning for Convolutional Models with a Study Case on Optical Coherence Tomography Image Classification
by: Rivera, Mariano, et al.
Published: (2025)
by: Rivera, Mariano, et al.
Published: (2025)
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
by: Toker, Michael, et al.
Published: (2024)
by: Toker, Michael, et al.
Published: (2024)
StereoCrafter: Diffusion-based Generation of Long and High-fidelity Stereoscopic 3D from Monocular Videos
by: Zhao, Sijie, et al.
Published: (2024)
by: Zhao, Sijie, et al.
Published: (2024)
Synthetic Image Generation in Cyber Influence Operations: An Emergent Threat?
by: Mathys, Melanie, et al.
Published: (2024)
by: Mathys, Melanie, et al.
Published: (2024)
Designing UNICORN: a Unified Benchmark for Imaging in Computational Pathology, Radiology, and Natural Language
by: Stegeman, Michelle, et al.
Published: (2026)
by: Stegeman, Michelle, et al.
Published: (2026)
Data-Augmented Multimodal Feature Fusion for Multiclass Visual Recognition of Oral Cancer Lesions
by: Naoum, Joy, et al.
Published: (2025)
by: Naoum, Joy, et al.
Published: (2025)
Threats and Opportunities in AI-generated Images for Armed Forces
by: Meier, Raphael
Published: (2025)
by: Meier, Raphael
Published: (2025)
Robust Self-calibration of Focal Lengths from the Fundamental Matrix
by: Kocur, Viktor, et al.
Published: (2023)
by: Kocur, Viktor, et al.
Published: (2023)
High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models
by: He, Mengqi, et al.
Published: (2025)
by: He, Mengqi, et al.
Published: (2025)
Detecting Effects of AI-Mediated Communication on Language Complexity and Sentiment
by: Sussman, Kristen, et al.
Published: (2025)
by: Sussman, Kristen, et al.
Published: (2025)
Unified Local and Global Attention Interaction Modeling for Vision Transformers
by: Nguyen, Tan, et al.
Published: (2024)
by: Nguyen, Tan, et al.
Published: (2024)
Efficient Diffusion Training through Parallelization with Truncated Karhunen-Loève Expansion
by: Ren, Yumeng, et al.
Published: (2025)
by: Ren, Yumeng, et al.
Published: (2025)
Event-based Solutions for Human-centered Applications: A Comprehensive Review
by: Adra, Mira, et al.
Published: (2025)
by: Adra, Mira, et al.
Published: (2025)
Grandes modelos de lenguaje: de la predicción de palabras a la comprensión?
by: Gómez-Rodríguez, Carlos
Published: (2025)
by: Gómez-Rodríguez, Carlos
Published: (2025)
Making High-Level AI Design Decisions Explicit Using a Binary Stream System-Designation Approach
by: Mossbridge, Julia
Published: (2024)
by: Mossbridge, Julia
Published: (2024)
Doodle Your Keypoints: Sketch-Based Few-Shot Keypoint Detection
by: Maity, Subhajit, et al.
Published: (2025)
by: Maity, Subhajit, et al.
Published: (2025)
PCRI: Measuring Context Robustness in Multimodal Models for Enterprise Applications
by: Patel, Hitesh Laxmichand, et al.
Published: (2025)
by: Patel, Hitesh Laxmichand, et al.
Published: (2025)
A Review of Pseudo-Labeling for Computer Vision
by: Kage, Patrick, et al.
Published: (2024)
by: Kage, Patrick, et al.
Published: (2024)
Sparse vs Contiguous Adversarial Pixel Perturbations in Multimodal Models: An Empirical Analysis
by: Botocan, Cristian-Alexandru, et al.
Published: (2024)
by: Botocan, Cristian-Alexandru, et al.
Published: (2024)
GeoPos: A Minimal Positional Encoding for Enhanced Fine-Grained Details in Image Synthesis Using Convolutional Neural Networks
by: Hosseini, Mehran, et al.
Published: (2024)
by: Hosseini, Mehran, et al.
Published: (2024)
Seeing The Words: Evaluating AI-generated Biblical Art
by: Makimei, Hidde, et al.
Published: (2025)
by: Makimei, Hidde, et al.
Published: (2025)
A Big Data Approach to Understand Sub-national Determinants of FDI in Africa
by: Colladon, A. Fronzetti, et al.
Published: (2024)
by: Colladon, A. Fronzetti, et al.
Published: (2024)
FLOWING: Implicit Neural Flows for Structure-Preserving Morphing
by: Bizzi, Arthur, et al.
Published: (2025)
by: Bizzi, Arthur, et al.
Published: (2025)
PlacidDreamer: Advancing Harmony in Text-to-3D Generation
by: Huang, Shuo, et al.
Published: (2024)
by: Huang, Shuo, et al.
Published: (2024)
Stereo Vision Based Robot for Remote Monitoring with VR Support
by: S., Mohamed Fazil M., et al.
Published: (2024)
by: S., Mohamed Fazil M., et al.
Published: (2024)
From Images to Decisions: Assistive Computer Vision for Non-Metallic Content Estimation in Scrap Metal
by: Storonkin, Daniil, et al.
Published: (2026)
by: Storonkin, Daniil, et al.
Published: (2026)
Textured-GS: Gaussian Splatting with Spatially Defined Color and Opacity
by: Huang, Zhentao, et al.
Published: (2024)
by: Huang, Zhentao, et al.
Published: (2024)
Robust automatic brain vessel segmentation in 3D CTA scans using dynamic 4D-CTA data
by: Ceballos-Arroyo, Alberto Mario, et al.
Published: (2026)
by: Ceballos-Arroyo, Alberto Mario, et al.
Published: (2026)
View-Consistent 3D Scene Editing via Dual-Path Structural Correspondense and Semantic Continuity
by: Li, Pufan, et al.
Published: (2026)
by: Li, Pufan, et al.
Published: (2026)
BEMTrace: Visualization-driven approach for deriving Building Energy Models from BIM
by: Walch, Andreas, et al.
Published: (2024)
by: Walch, Andreas, et al.
Published: (2024)
Dream Content Discovery from Reddit with an Unsupervised Mixed-Method Approach
by: Das, Anubhab, et al.
Published: (2023)
by: Das, Anubhab, et al.
Published: (2023)
ANCHOR: LLM-driven Subject Conditioning for Text-to-Image Synthesis
by: Ramakrishnan, Aashish Anantha, et al.
Published: (2024)
by: Ramakrishnan, Aashish Anantha, et al.
Published: (2024)
STRICT: Stress Test of Rendering Images Containing Text
by: Zhang, Tianyu, et al.
Published: (2025)
by: Zhang, Tianyu, et al.
Published: (2025)
Once-For-All: A Train-Once and Select-Anytime Framework for Multimodal Instruction Tuning
by: Dong, Mingkang, et al.
Published: (2026)
by: Dong, Mingkang, et al.
Published: (2026)
The Power of Absence: Thinking with Archival Theory in Algorithmic Design
by: Sherman, Jihan, et al.
Published: (2024)
by: Sherman, Jihan, et al.
Published: (2024)
Similar Items
-
AzSLD: Azerbaijani Sign Language Dataset for Fingerspelling, Word, and Sentence Translation with Baseline Software
by: Alishzade, Nigar, et al.
Published: (2024) -
Proto-FG3D: Prototype-based Interpretable Fine-Grained 3D Shape Classification
by: Ma, Shuxian, et al.
Published: (2025) -
CerberusDet: Unified Multi-Dataset Object Detection
by: Tolstykh, Irina, et al.
Published: (2024) -
Eye-gaze Guided Multi-modal Alignment for Medical Representation Learning
by: Ma, Chong, et al.
Published: (2024) -
Synthetic Photography Detection: A Visual Guidance for Identifying Synthetic Images Created by AI
by: Mathys, Melanie, et al.
Published: (2024)