Saved in:
| Main Authors: | Basioti, Kalliopi, Sahu, Pritish, Liu, Qingze Tony, Xu, Zihao, Wang, Hao, Pavlovic, Vladimir |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2503.23598 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Box2Flow: Instance-based Action Flow Graphs from Videos
by: Li, Jiatong, et al.
Published: (2024)
by: Li, Jiatong, et al.
Published: (2024)
CIC-BART-SSA: Controllable Image Captioning with Structured Semantic Augmentation
by: Basioti, Kalliopi, et al.
Published: (2024)
by: Basioti, Kalliopi, et al.
Published: (2024)
CASIM: Composite Aware Semantic Injection for Text to Motion Generation
by: Chang, Che-Jui, et al.
Published: (2025)
by: Chang, Che-Jui, et al.
Published: (2025)
Hallucinatory Image Tokens: A Training-free EAZY Approach on Detecting and Mitigating Object Hallucinations in LVLMs
by: Che, Liwei, et al.
Published: (2025)
by: Che, Liwei, et al.
Published: (2025)
HCVP: Leveraging Hierarchical Contrastive Visual Prompt for Domain Generalization
by: Zhou, Guanglin, et al.
Published: (2024)
by: Zhou, Guanglin, et al.
Published: (2024)
TrajDiffuse: A Conditional Diffusion Model for Environment-Aware Trajectory Prediction
by: Qingze, et al.
Published: (2024)
by: Qingze, et al.
Published: (2024)
CODA: Repurposing Continuous VAEs for Discrete Tokenization
by: Liu, Zeyu, et al.
Published: (2025)
by: Liu, Zeyu, et al.
Published: (2025)
Hierarchical Concept-to-Appearance Guidance for Multi-Subject Image Generation
by: Xu, Yijia, et al.
Published: (2026)
by: Xu, Yijia, et al.
Published: (2026)
GenEscape: Hierarchical Multi-Agent Generation of Escape Room Puzzles
by: Shan, Mengyi, et al.
Published: (2025)
by: Shan, Mengyi, et al.
Published: (2025)
Enhancing Consistency Models for Multi-Agent Trajectory Prediction
by: Mrdovic, Alen, et al.
Published: (2026)
by: Mrdovic, Alen, et al.
Published: (2026)
GenIR: Generative Visual Feedback for Mental Image Retrieval
by: Yang, Diji, et al.
Published: (2025)
by: Yang, Diji, et al.
Published: (2025)
Learning Energy-based Variational Latent Prior for VAEs
by: Dutta, Debottam, et al.
Published: (2025)
by: Dutta, Debottam, et al.
Published: (2025)
TeaserGen: Generating Teasers for Long Documentaries
by: Xu, Weihan, et al.
Published: (2024)
by: Xu, Weihan, et al.
Published: (2024)
Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models
by: Che, Liwei, et al.
Published: (2026)
by: Che, Liwei, et al.
Published: (2026)
GameGen-X: Interactive Open-world Game Video Generation
by: Che, Haoxuan, et al.
Published: (2024)
by: Che, Haoxuan, et al.
Published: (2024)
Benchmarking Content-Based Puzzle Solvers on Corrupted Jigsaw Puzzles
by: Dirauf, Richard, et al.
Published: (2025)
by: Dirauf, Richard, et al.
Published: (2025)
Adversarial robustness of VAEs through the lens of local geometry
by: Khan, Asif, et al.
Published: (2022)
by: Khan, Asif, et al.
Published: (2022)
GenArena: How Can We Achieve Human-Aligned Evaluation for Visual Generation Tasks?
by: Li, Ruihang, et al.
Published: (2026)
by: Li, Ruihang, et al.
Published: (2026)
VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL
by: Feng, Yichen, et al.
Published: (2025)
by: Feng, Yichen, et al.
Published: (2025)
PuzzleBench: A Fully Dynamic Evaluation Framework for Large Multimodal Models on Puzzle Solving
by: Zhang, Zeyu, et al.
Published: (2025)
by: Zhang, Zeyu, et al.
Published: (2025)
Action Quality Assessment via Hierarchical Pose-guided Multi-stage Contrastive Regression
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
Feature Identification for Hierarchical Contrastive Learning
by: Ott, Julius, et al.
Published: (2025)
by: Ott, Julius, et al.
Published: (2025)
VP-LLM: Text-Driven 3D Volume Completion with Large Language Models through Patchification
by: Liu, Jianmeng, et al.
Published: (2024)
by: Liu, Jianmeng, et al.
Published: (2024)
AdverX-Ray: Ensuring X-Ray Integrity Through Frequency-Sensitive Adversarial VAEs
by: Caetano, Francisco, et al.
Published: (2025)
by: Caetano, Francisco, et al.
Published: (2025)
Are Language Models Puzzle Prodigies? Algorithmic Puzzles Unveil Serious Challenges in Multimodal Reasoning
by: Ghosal, Deepanway, et al.
Published: (2024)
by: Ghosal, Deepanway, et al.
Published: (2024)
Gen4Gen: Generative Data Pipeline for Generative Multi-Concept Composition
by: Yeh, Chun-Hsiao, et al.
Published: (2024)
by: Yeh, Chun-Hsiao, et al.
Published: (2024)
ConVQG: Contrastive Visual Question Generation with Multimodal Guidance
by: Mi, Li, et al.
Published: (2024)
by: Mi, Li, et al.
Published: (2024)
OmniGen: Unified Image Generation
by: Xiao, Shitao, et al.
Published: (2024)
by: Xiao, Shitao, et al.
Published: (2024)
Understanding Model Calibration -- A gentle introduction and visual exploration of calibration and the expected calibration error (ECE)
by: Pavlovic, Maja
Published: (2025)
by: Pavlovic, Maja
Published: (2025)
Eye-Q: A Multilingual Benchmark for Visual Word Puzzle Solving and Image-to-Phrase Reasoning
by: Najar, Ali, et al.
Published: (2026)
by: Najar, Ali, et al.
Published: (2026)
Deep Incomplete Multi-view Learning via Cyclic Permutation of VAEs
by: Gao, Xin, et al.
Published: (2025)
by: Gao, Xin, et al.
Published: (2025)
PuzzleGPT: Emulating Human Puzzle-Solving Ability for Time and Location Prediction
by: Ayyubi, Hammad, et al.
Published: (2025)
by: Ayyubi, Hammad, et al.
Published: (2025)
Hyperbolic Contrastive Learning for Hierarchical 3D Point Cloud Embedding
by: Liu, Yingjie, et al.
Published: (2025)
by: Liu, Yingjie, et al.
Published: (2025)
Foundation VAEs for 3D CT Reconstruction, Augmentation, and Generation
by: Chen, Qi, et al.
Published: (2026)
by: Chen, Qi, et al.
Published: (2026)
GenShield: Unified Detection and Artifact Correction for AI-Generated Images
by: Xu, Zhipei, et al.
Published: (2026)
by: Xu, Zhipei, et al.
Published: (2026)
Gen-AFFECT: Generation of Avatar Fine-grained Facial Expressions with Consistent identiTy
by: Yu, Hao, et al.
Published: (2025)
by: Yu, Hao, et al.
Published: (2025)
Pyramid Coder: Hierarchical Code Generator for Compositional Visual Question Answering
by: Shen, Ruoyue, et al.
Published: (2024)
by: Shen, Ruoyue, et al.
Published: (2024)
Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding
by: Wang, Haibo, et al.
Published: (2026)
by: Wang, Haibo, et al.
Published: (2026)
GenFusion: Closing the Loop between Reconstruction and Generation via Videos
by: Wu, Sibo, et al.
Published: (2025)
by: Wu, Sibo, et al.
Published: (2025)
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
by: Ge, Yuyao, et al.
Published: (2025)
by: Ge, Yuyao, et al.
Published: (2025)
Similar Items
-
Box2Flow: Instance-based Action Flow Graphs from Videos
by: Li, Jiatong, et al.
Published: (2024) -
CIC-BART-SSA: Controllable Image Captioning with Structured Semantic Augmentation
by: Basioti, Kalliopi, et al.
Published: (2024) -
CASIM: Composite Aware Semantic Injection for Text to Motion Generation
by: Chang, Che-Jui, et al.
Published: (2025) -
Hallucinatory Image Tokens: A Training-free EAZY Approach on Detecting and Mitigating Object Hallucinations in LVLMs
by: Che, Liwei, et al.
Published: (2025) -
HCVP: Leveraging Hierarchical Contrastive Visual Prompt for Domain Generalization
by: Zhou, Guanglin, et al.
Published: (2024)