Gespeichert in:
| Hauptverfasser: | Sushko, Peter, Bharadwaj, Ayana, Lim, Zhi Yang, Ilin, Vasily, Caffee, Ben, Chen, Dongping, Salehi, Mohammadreza, Hsieh, Cheng-Yu, Krishna, Ranjay |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2502.03629 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MultiRef: Controllable Image Generation with Multiple Visual References
von: Chen, Ruoxi, et al.
Veröffentlicht: (2025)
von: Chen, Ruoxi, et al.
Veröffentlicht: (2025)
VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition
von: Yadav, Tanush, et al.
Veröffentlicht: (2026)
von: Yadav, Tanush, et al.
Veröffentlicht: (2026)
Score-based deterministic density sampling
von: Ilin, Vasily, et al.
Veröffentlicht: (2025)
von: Ilin, Vasily, et al.
Veröffentlicht: (2025)
Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency
von: Wang, Chenlong, et al.
Veröffentlicht: (2025)
von: Wang, Chenlong, et al.
Veröffentlicht: (2025)
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
von: Bigverdi, Mahtab, et al.
Veröffentlicht: (2024)
von: Bigverdi, Mahtab, et al.
Veröffentlicht: (2024)
The Hard Positive Truth about Vision-Language Compositionality
von: Kamath, Amita, et al.
Veröffentlicht: (2024)
von: Kamath, Amita, et al.
Veröffentlicht: (2024)
ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2024)
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2024)
One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory
von: Zheng, Chenhao, et al.
Veröffentlicht: (2025)
von: Zheng, Chenhao, et al.
Veröffentlicht: (2025)
The Chronicles of RiDiC: Generating Datasets with Controlled Popularity Distribution for Long-form Factuality Evaluation
von: Braslavski, Pavel, et al.
Veröffentlicht: (2026)
von: Braslavski, Pavel, et al.
Veröffentlicht: (2026)
Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
Is C4 Dataset Optimal for Pruning? An Investigation of Calibration Data for LLM Pruning
von: Bandari, Abhinav, et al.
Veröffentlicht: (2024)
von: Bandari, Abhinav, et al.
Veröffentlicht: (2024)
Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2024)
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2024)
The Unmet Promise of Synthetic Training Images: Using Retrieved Real Images Performs Better
von: Geng, Scott, et al.
Veröffentlicht: (2024)
von: Geng, Scott, et al.
Veröffentlicht: (2024)
Reinforced Visual Perception with Tools
von: Zhou, Zetong, et al.
Veröffentlicht: (2025)
von: Zhou, Zetong, et al.
Veröffentlicht: (2025)
Seeking and Updating with Live Visual Knowledge
von: Fu, Mingyang, et al.
Veröffentlicht: (2025)
von: Fu, Mingyang, et al.
Veröffentlicht: (2025)
OmniView: An All-Seeing Diffusion Model for 3D and 4D View Synthesis
von: Fan, Xiang, et al.
Veröffentlicht: (2025)
von: Fan, Xiang, et al.
Veröffentlicht: (2025)
Iterated Learning Improves Compositionality in Large Vision-Language Models
von: Zheng, Chenhao, et al.
Veröffentlicht: (2024)
von: Zheng, Chenhao, et al.
Veröffentlicht: (2024)
Semantic and Expressive Variation in Image Captions Across Languages
von: Ye, Andre, et al.
Veröffentlicht: (2023)
von: Ye, Andre, et al.
Veröffentlicht: (2023)
Agonistic Image Generation: Unsettling the Hegemony of Intention
von: Shaw, Andrew, et al.
Veröffentlicht: (2025)
von: Shaw, Andrew, et al.
Veröffentlicht: (2025)
FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations
von: Hsieh, Cheng-Yu, et al.
Veröffentlicht: (2025)
von: Hsieh, Cheng-Yu, et al.
Veröffentlicht: (2025)
Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?
von: Zhang, Tianyi, et al.
Veröffentlicht: (2026)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2026)
PulseReddit: A Novel Reddit Dataset for Benchmarking MAS in High-Frequency Cryptocurrency Trading
von: Han, Qiuhan, et al.
Veröffentlicht: (2025)
von: Han, Qiuhan, et al.
Veröffentlicht: (2025)
SPARO: Selective Attention for Robust and Compositional Transformer Encodings for Vision
von: Vani, Ankit, et al.
Veröffentlicht: (2024)
von: Vani, Ankit, et al.
Veröffentlicht: (2024)
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
von: Clark, Christopher, et al.
Veröffentlicht: (2026)
von: Clark, Christopher, et al.
Veröffentlicht: (2026)
EditSleuth: A Dataset of Grounded Reasoning Chains for Image-Edit Forensics
von: Nguyen, Van-Loc, et al.
Veröffentlicht: (2026)
von: Nguyen, Van-Loc, et al.
Veröffentlicht: (2026)
TECCI: Tricky Edits of Collected and Curated Images
von: Agrawal, Aishwarya, et al.
Veröffentlicht: (2026)
von: Agrawal, Aishwarya, et al.
Veröffentlicht: (2026)
GalaxyEdit: Large-Scale Image Editing Dataset with Enhanced Diffusion Adapter
von: Bala, Aniruddha, et al.
Veröffentlicht: (2024)
von: Bala, Aniruddha, et al.
Veröffentlicht: (2024)
From Reddit to Generative AI: Evaluating Large Language Models for Anxiety Support Fine-tuned on Social Media Data
von: Kursuncu, Ugur, et al.
Veröffentlicht: (2025)
von: Kursuncu, Ugur, et al.
Veröffentlicht: (2025)
Videoshop: Localized Semantic Video Editing with Noise-Extrapolated Diffusion Inversion
von: Fan, Xiang, et al.
Veröffentlicht: (2024)
von: Fan, Xiang, et al.
Veröffentlicht: (2024)
Dynamic Template Selection for Output Token Generation Optimization: MLP-Based and Transformer Approaches
von: Yadavalli, Bharadwaj
Veröffentlicht: (2025)
von: Yadavalli, Bharadwaj
Veröffentlicht: (2025)
m&m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks
von: Ma, Zixian, et al.
Veröffentlicht: (2024)
von: Ma, Zixian, et al.
Veröffentlicht: (2024)
FunEditor: Achieving Complex Image Edits via Function Aggregation with Diffusion Models
von: Samadi, Mohammadreza, et al.
Veröffentlicht: (2024)
von: Samadi, Mohammadreza, et al.
Veröffentlicht: (2024)
ImgEdit: A Unified Image Editing Dataset and Benchmark
von: Ye, Yang, et al.
Veröffentlicht: (2025)
von: Ye, Yang, et al.
Veröffentlicht: (2025)
DiT4Edit: Diffusion Transformer for Image Editing
von: Feng, Kunyu, et al.
Veröffentlicht: (2024)
von: Feng, Kunyu, et al.
Veröffentlicht: (2024)
TWeddit : A Dataset of Triggering Stories Predominantly Shared by Women on Reddit
von: Bandela, Shirlene Rose, et al.
Veröffentlicht: (2026)
von: Bandela, Shirlene Rose, et al.
Veröffentlicht: (2026)
In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer
von: Zhang, Zechuan, et al.
Veröffentlicht: (2025)
von: Zhang, Zechuan, et al.
Veröffentlicht: (2025)
Throwaway Accounts and Moderation on Reddit
von: Guo, Cheng, et al.
Veröffentlicht: (2025)
von: Guo, Cheng, et al.
Veröffentlicht: (2025)
ImageInWords: Unlocking Hyper-Detailed Image Descriptions
von: Garg, Roopal, et al.
Veröffentlicht: (2024)
von: Garg, Roopal, et al.
Veröffentlicht: (2024)
The Moral Foundations Reddit Corpus
von: Trager, Jackson, et al.
Veröffentlicht: (2022)
von: Trager, Jackson, et al.
Veröffentlicht: (2022)
GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
von: Kamath, Amita, et al.
Veröffentlicht: (2025)
von: Kamath, Amita, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MultiRef: Controllable Image Generation with Multiple Visual References
von: Chen, Ruoxi, et al.
Veröffentlicht: (2025) -
VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition
von: Yadav, Tanush, et al.
Veröffentlicht: (2026) -
Score-based deterministic density sampling
von: Ilin, Vasily, et al.
Veröffentlicht: (2025) -
Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency
von: Wang, Chenlong, et al.
Veröffentlicht: (2025) -
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
von: Bigverdi, Mahtab, et al.
Veröffentlicht: (2024)