Saved in:
| Main Authors: | Tang, Kexian, Gao, Junyao, Zeng, Yanhong, Duan, Haodong, Sun, Yanan, Xing, Zhening, Liu, Wenran, Lyu, Kaifeng, Chen, Kai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2503.19990 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CharacterShot: Controllable and Consistent 4D Character Animation
by: Gao, Junyao, et al.
Published: (2025)
by: Gao, Junyao, et al.
Published: (2025)
MotionShot: Adaptive Motion Transfer across Arbitrary Objects for Text-to-Video Generation
by: Liu, Yanchen, et al.
Published: (2025)
by: Liu, Yanchen, et al.
Published: (2025)
SPA: A Simple but Tough-to-Beat Baseline for Knowledge Injection
by: Tang, Kexian, et al.
Published: (2026)
by: Tang, Kexian, et al.
Published: (2026)
FaceShot: Bring Any Character into Life
by: Gao, Junyao, et al.
Published: (2025)
by: Gao, Junyao, et al.
Published: (2025)
StyleShot: A Snapshot on Any Style
by: Gao, Junyao, et al.
Published: (2024)
by: Gao, Junyao, et al.
Published: (2024)
PIA: Your Personalized Image Animator via Plug-and-Play Modules in Text-to-Image Models
by: Zhang, Yiming, et al.
Published: (2023)
by: Zhang, Yiming, et al.
Published: (2023)
A Task is Worth One Word: Learning with Task Prompts for High-Quality Versatile Image Inpainting
by: Zhuang, Junhao, et al.
Published: (2023)
by: Zhuang, Junhao, et al.
Published: (2023)
Adam Reduces a Unique Form of Sharpness: Theoretical Insights Near the Minimizer Manifold
by: Li, Xinghan, et al.
Published: (2025)
by: Li, Xinghan, et al.
Published: (2025)
Live2Diff: Live Stream Translation via Uni-directional Attention in Video Diffusion Models
by: Xing, Zhening, et al.
Published: (2024)
by: Xing, Zhening, et al.
Published: (2024)
FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
How Good are Foundation Models in Step-by-Step Embodied Reasoning?
by: Dissanayake, Dinura, et al.
Published: (2025)
by: Dissanayake, Dinura, et al.
Published: (2025)
MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI
by: Yao, Huanjin, et al.
Published: (2025)
by: Yao, Huanjin, et al.
Published: (2025)
Redundancy Principles for MLLMs Benchmarks
by: Zhang, Zicheng, et al.
Published: (2025)
by: Zhang, Zicheng, et al.
Published: (2025)
Pencil Puzzle Bench: A Benchmark for Multi-Step Verifiable Reasoning
by: Waugh, Justin
Published: (2026)
by: Waugh, Justin
Published: (2026)
Fine-tuning MLLMs Without Forgetting Is Easier Than You Think
by: Li, He, et al.
Published: (2026)
by: Li, He, et al.
Published: (2026)
Shift is Good: Mismatched Data Mixing Improves Test Performance
by: Medvedev, Marko, et al.
Published: (2025)
by: Medvedev, Marko, et al.
Published: (2025)
How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM Pretraining
by: Luo, Kairong, et al.
Published: (2025)
by: Luo, Kairong, et al.
Published: (2025)
A Multi-Power Law for Loss Curve Prediction Across Learning Rate Schedules
by: Luo, Kairong, et al.
Published: (2025)
by: Luo, Kairong, et al.
Published: (2025)
Larger Datasets Can Be Repeated More: A Theoretical Analysis of Multi-Epoch Scaling in Linear Regression
by: Yan, Tingkai, et al.
Published: (2025)
by: Yan, Tingkai, et al.
Published: (2025)
TangramPuzzle: Evaluating Multimodal Large Language Models with Compositional Spatial Reasoning
by: Liu, Daixian, et al.
Published: (2026)
by: Liu, Daixian, et al.
Published: (2026)
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
by: Ouyang, Kun, et al.
Published: (2025)
by: Ouyang, Kun, et al.
Published: (2025)
Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing
by: Zhao, Xiangyu, et al.
Published: (2025)
by: Zhao, Xiangyu, et al.
Published: (2025)
Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?
by: Tyagi, Nemika, et al.
Published: (2024)
by: Tyagi, Nemika, et al.
Published: (2024)
Affordance Benchmark for MLLMs
by: Wang, Junying, et al.
Published: (2025)
by: Wang, Junying, et al.
Published: (2025)
How Far Are We from Optimal Reasoning Efficiency?
by: Gao, Jiaxuan, et al.
Published: (2025)
by: Gao, Jiaxuan, et al.
Published: (2025)
Video-MSR: Benchmarking Multi-hop Spatial Reasoning Capabilities of MLLMs
by: Zhu, Rui, et al.
Published: (2026)
by: Zhu, Rui, et al.
Published: (2026)
Human-Agent Collaborative Paper-to-Page Crafting for Under $0.1
by: Ma, Qianli, et al.
Published: (2025)
by: Ma, Qianli, et al.
Published: (2025)
SpatialTree: How Spatial Abilities Branch Out in MLLMs
by: Xiao, Yuxi, et al.
Published: (2025)
by: Xiao, Yuxi, et al.
Published: (2025)
LEGO: Spatial Accelerator Generation and Optimization for Tensor Applications
by: Lin, Yujun, et al.
Published: (2025)
by: Lin, Yujun, et al.
Published: (2025)
Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs
by: Sun, Yanpeng, et al.
Published: (2025)
by: Sun, Yanpeng, et al.
Published: (2025)
What Defines Good Reasoning in LLMs? Dissecting Reasoning Steps with Multi-Aspect Evaluation
by: Do, Heejin, et al.
Published: (2025)
by: Do, Heejin, et al.
Published: (2025)
Spatial Preference Rewarding for MLLMs Spatial Understanding
by: Qiu, Han, et al.
Published: (2025)
by: Qiu, Han, et al.
Published: (2025)
M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
by: AI, Inclusion, et al.
Published: (2025)
by: AI, Inclusion, et al.
Published: (2025)
PuzzlePlex: Benchmarking Foundation Models on Reasoning and Planning with Puzzles
by: Long, Yitao, et al.
Published: (2025)
by: Long, Yitao, et al.
Published: (2025)
MonetGPT: Solving Puzzles Enhances MLLMs' Image Retouching Skills
by: Dutt, Niladri Shekhar, et al.
Published: (2025)
by: Dutt, Niladri Shekhar, et al.
Published: (2025)
Dataset Distillers Are Good Label Denoisers In the Wild
by: Cheng, Lechao, et al.
Published: (2024)
by: Cheng, Lechao, et al.
Published: (2024)
GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs
by: Zhu, Xiaorong, et al.
Published: (2025)
by: Zhu, Xiaorong, et al.
Published: (2025)
Scaling Spatial Reasoning in MLLMs through Programmatic Data Synthesis
by: Helu, Zhi, et al.
Published: (2025)
by: Helu, Zhi, et al.
Published: (2025)
Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs
by: Anand, Dhruv, et al.
Published: (2025)
by: Anand, Dhruv, et al.
Published: (2025)
Displaced Fermionic Gaussian States and their Classical Simulation
by: Lyu, Xingjian, et al.
Published: (2024)
by: Lyu, Xingjian, et al.
Published: (2024)
Similar Items
-
CharacterShot: Controllable and Consistent 4D Character Animation
by: Gao, Junyao, et al.
Published: (2025) -
MotionShot: Adaptive Motion Transfer across Arbitrary Objects for Text-to-Video Generation
by: Liu, Yanchen, et al.
Published: (2025) -
SPA: A Simple but Tough-to-Beat Baseline for Knowledge Injection
by: Tang, Kexian, et al.
Published: (2026) -
FaceShot: Bring Any Character into Life
by: Gao, Junyao, et al.
Published: (2025) -
StyleShot: A Snapshot on Any Style
by: Gao, Junyao, et al.
Published: (2024)