Salvato in:
| Autore principale: | Fixelle, Joshua |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2504.08710 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
More than the Sum: Panorama-Language Models for Adverse Omni-Scenes
di: Fan, Weijia, et al.
Pubblicazione: (2026)
di: Fan, Weijia, et al.
Pubblicazione: (2026)
More than a Moment: Towards Coherent Sequences of Audio Descriptions
di: Khandelwal, Eshika, et al.
Pubblicazione: (2025)
di: Khandelwal, Eshika, et al.
Pubblicazione: (2025)
Vision Transformers Need More Than Registers
di: Shi, Cheng, et al.
Pubblicazione: (2026)
di: Shi, Cheng, et al.
Pubblicazione: (2026)
Opinion: Learning Intuitive Physics May Require More than Visual Data
di: Su, Ellen, et al.
Pubblicazione: (2025)
di: Su, Ellen, et al.
Pubblicazione: (2025)
More than the Sum of Its Parts: Ensembling Backbone Networks for Few-Shot Segmentation
di: Catalano, Nico, et al.
Pubblicazione: (2024)
di: Catalano, Nico, et al.
Pubblicazione: (2024)
Nearly Solved? Robust Deepfake Detection Requires More than Visual Forensics
di: Levy, Guy, et al.
Pubblicazione: (2024)
di: Levy, Guy, et al.
Pubblicazione: (2024)
More than Segmentation: Benchmarking SAM 3 for Segmentation, 3D Perception, and Reconstruction in Robotic Surgery
di: Dong, Wenzhen, et al.
Pubblicazione: (2025)
di: Dong, Wenzhen, et al.
Pubblicazione: (2025)
There is More to Attention: Statistical Filtering Enhances Explanations in Vision Transformers
di: Ayyar, Meghna P, et al.
Pubblicazione: (2025)
di: Ayyar, Meghna P, et al.
Pubblicazione: (2025)
More than One Step at a Time: Designing Procedural Feedback for Non-visual Makeup Routines
di: Li, Franklin Mingzhe, et al.
Pubblicazione: (2025)
di: Li, Franklin Mingzhe, et al.
Pubblicazione: (2025)
Nodes Are Early, Edges Are Late: Probing Diagram Representations in Large Vision-Language Models
di: Yoshida, Haruto, et al.
Pubblicazione: (2026)
di: Yoshida, Haruto, et al.
Pubblicazione: (2026)
More than Memes: A Multimodal Topic Modeling Approach to Conspiracy Theories on Telegram
di: Steffen, Elisabeth
Pubblicazione: (2024)
di: Steffen, Elisabeth
Pubblicazione: (2024)
AdaNCA: Neural Cellular Automata As Adaptors For More Robust Vision Transformer
di: Xu, Yitao, et al.
Pubblicazione: (2024)
di: Xu, Yitao, et al.
Pubblicazione: (2024)
SVD-ViT: Does SVD Make Vision Transformers Attend More to the Foreground?
di: Murata, Haruhiko, et al.
Pubblicazione: (2026)
di: Murata, Haruhiko, et al.
Pubblicazione: (2026)
More Images, More Problems? A Controlled Analysis of VLM Failure Modes
di: Das, Anurag, et al.
Pubblicazione: (2026)
di: Das, Anurag, et al.
Pubblicazione: (2026)
Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition
di: Surikuchi, Aditya K, et al.
Pubblicazione: (2024)
di: Surikuchi, Aditya K, et al.
Pubblicazione: (2024)
Representation Alignment for Just Image Transformers is not Easier than You Think
di: Shin, Jaeyo, et al.
Pubblicazione: (2026)
di: Shin, Jaeyo, et al.
Pubblicazione: (2026)
From Edges to Depth: Probing the Spatial Hierarchy in Vision Transformers
di: Sanghavi, Jainum
Pubblicazione: (2026)
di: Sanghavi, Jainum
Pubblicazione: (2026)
Towards Efficient Vision-Language Tuning: More Information Density, More Generalizability
di: Hao, Tianxiang, et al.
Pubblicazione: (2023)
di: Hao, Tianxiang, et al.
Pubblicazione: (2023)
More Clear, More Flexible, More Precise: A Comprehensive Oriented Object Detection benchmark for UAV
di: Ye, Kai, et al.
Pubblicazione: (2025)
di: Ye, Kai, et al.
Pubblicazione: (2025)
Adapted Center and Scale Prediction: More Stable and More Accurate
di: Wang, Wenhao, et al.
Pubblicazione: (2020)
di: Wang, Wenhao, et al.
Pubblicazione: (2020)
Do More Details Always Introduce More Hallucinations in LVLM-based Image Captioning?
di: Feng, Mingqian, et al.
Pubblicazione: (2024)
di: Feng, Mingqian, et al.
Pubblicazione: (2024)
Alignment and Adversarial Robustness: Are More Human-Like Models More Secure?
di: Hoak, Blaine, et al.
Pubblicazione: (2025)
di: Hoak, Blaine, et al.
Pubblicazione: (2025)
Leaner Transformers: More Heads, Less Depth
di: Saratchandran, Hemanth, et al.
Pubblicazione: (2025)
di: Saratchandran, Hemanth, et al.
Pubblicazione: (2025)
Less is More: Skim Transformer for Light Field Image Super-resolution
di: Hu, Zeke Zexi, et al.
Pubblicazione: (2024)
di: Hu, Zeke Zexi, et al.
Pubblicazione: (2024)
The More You See in 2D, the More You Perceive in 3D
di: Han, Xinyang, et al.
Pubblicazione: (2024)
di: Han, Xinyang, et al.
Pubblicazione: (2024)
Larger than memory image processing
di: Sporring, Jon, et al.
Pubblicazione: (2026)
di: Sporring, Jon, et al.
Pubblicazione: (2026)
Vision-Language Models Generate More Homogeneous Stories for Phenotypically Black Individuals
di: Lee, Messi H. J., et al.
Pubblicazione: (2024)
di: Lee, Messi H. J., et al.
Pubblicazione: (2024)
ForensicZip: More Tokens are Better but Not Necessary in Forensic Vision-Language Models
di: Lai, Yingxin, et al.
Pubblicazione: (2026)
di: Lai, Yingxin, et al.
Pubblicazione: (2026)
Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors
di: Wang, Xiangchen, et al.
Pubblicazione: (2025)
di: Wang, Xiangchen, et al.
Pubblicazione: (2025)
More Pictures Say More: Visual Intersection Network for Open Set Object Detection
di: Dong, Bingcheng, et al.
Pubblicazione: (2024)
di: Dong, Bingcheng, et al.
Pubblicazione: (2024)
Fine-Tuning Image-Conditional Diffusion Models is Easier than You Think
di: Garcia, Gonzalo Martin, et al.
Pubblicazione: (2024)
di: Garcia, Gonzalo Martin, et al.
Pubblicazione: (2024)
Pruning One More Token is Enough: Leveraging Latency-Workload Non-Linearities for Vision Transformers on the Edge
di: Eliopoulos, Nick John, et al.
Pubblicazione: (2024)
di: Eliopoulos, Nick John, et al.
Pubblicazione: (2024)
Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
di: Wang, Feng, et al.
Pubblicazione: (2025)
di: Wang, Feng, et al.
Pubblicazione: (2025)
Floating No More: Object-Ground Reconstruction from a Single Image
di: Man, Yunze, et al.
Pubblicazione: (2024)
di: Man, Yunze, et al.
Pubblicazione: (2024)
Reduce the Artifacts Bias for More Generalizable AI-Generated Image Detection
di: Li, Yiheng, et al.
Pubblicazione: (2026)
di: Li, Yiheng, et al.
Pubblicazione: (2026)
An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels
di: Nguyen, Duy-Kien, et al.
Pubblicazione: (2024)
di: Nguyen, Duy-Kien, et al.
Pubblicazione: (2024)
Less-to-More Generalization: Unlocking More Controllability by In-Context Generation
di: Wu, Shaojin, et al.
Pubblicazione: (2025)
di: Wu, Shaojin, et al.
Pubblicazione: (2025)
A Little More Like This: Text-to-Image Retrieval with Vision-Language Models Using Relevance Feedback
di: Khaertdinov, Bulat, et al.
Pubblicazione: (2025)
di: Khaertdinov, Bulat, et al.
Pubblicazione: (2025)
Towards More Accurate Personalized Image Generation: Addressing Overfitting and Evaluation Bias
di: Li, Mingxiao, et al.
Pubblicazione: (2025)
di: Li, Mingxiao, et al.
Pubblicazione: (2025)
How to Learn More? Exploring Kolmogorov-Arnold Networks for Hyperspectral Image Classification
di: Jamali, Ali, et al.
Pubblicazione: (2024)
di: Jamali, Ali, et al.
Pubblicazione: (2024)
Documenti analoghi
-
More than the Sum: Panorama-Language Models for Adverse Omni-Scenes
di: Fan, Weijia, et al.
Pubblicazione: (2026) -
More than a Moment: Towards Coherent Sequences of Audio Descriptions
di: Khandelwal, Eshika, et al.
Pubblicazione: (2025) -
Vision Transformers Need More Than Registers
di: Shi, Cheng, et al.
Pubblicazione: (2026) -
Opinion: Learning Intuitive Physics May Require More than Visual Data
di: Su, Ellen, et al.
Pubblicazione: (2025) -
More than the Sum of Its Parts: Ensembling Backbone Networks for Few-Shot Segmentation
di: Catalano, Nico, et al.
Pubblicazione: (2024)