FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations
Fuente:
arXiv
Salvato in:
| Autori principali: | Hsieh, Cheng-Yu, Vasu, Pavan Kumar Anasosalu, Faghri, Fartash, Vemulapalli, Raviteja, Li, Chun-Liang, Krishna, Ranjay, Tuzel, Oncel, Pouransari, Hadi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
di: Vasu, Pavan Kumar Anasosalu, et al.
Pubblicazione: (2023)
di: Vasu, Pavan Kumar Anasosalu, et al.
Pubblicazione: (2023)
CLIP with Quality Captions: A Strong Pretraining for Vision Tasks
di: Vasu, Pavan Kumar Anasosalu, et al.
Pubblicazione: (2024)
di: Vasu, Pavan Kumar Anasosalu, et al.
Pubblicazione: (2024)
SAM-CLIP: Merging Vision Foundation Models towards Semantic and Spatial Understanding
di: Wang, Haoxiang, et al.
Pubblicazione: (2023)
di: Wang, Haoxiang, et al.
Pubblicazione: (2023)
MobileCLIP2: Improving Multi-Modal Reinforced Training
di: Faghri, Fartash, et al.
Pubblicazione: (2025)
di: Faghri, Fartash, et al.
Pubblicazione: (2025)
Knowledge Transfer from Vision Foundation Models for Efficient Training of Small Task-specific Models
di: Vemulapalli, Raviteja, et al.
Pubblicazione: (2023)
di: Vemulapalli, Raviteja, et al.
Pubblicazione: (2023)
TiC-CLIP: Continual Training of CLIP Models
di: Garg, Saurabh, et al.
Pubblicazione: (2023)
di: Garg, Saurabh, et al.
Pubblicazione: (2023)
VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models
di: Vasu, Pavan Kumar Anasosalu, et al.
Pubblicazione: (2026)
di: Vasu, Pavan Kumar Anasosalu, et al.
Pubblicazione: (2026)
Proxy-FDA: Proxy-based Feature Distribution Alignment for Fine-tuning Vision Foundation Models without Forgetting
di: Huang, Chen, et al.
Pubblicazione: (2025)
di: Huang, Chen, et al.
Pubblicazione: (2025)
MUSCLE: A Model Update Strategy for Compatible LLM Evolution
di: Echterhoff, Jessica, et al.
Pubblicazione: (2024)
di: Echterhoff, Jessica, et al.
Pubblicazione: (2024)
FastVLM: Efficient Vision Encoding for Vision Language Models
di: Vasu, Pavan Kumar Anasosalu, et al.
Pubblicazione: (2024)
di: Vasu, Pavan Kumar Anasosalu, et al.
Pubblicazione: (2024)
TiC-LM: A Web-Scale Benchmark for Time-Continual LLM Pretraining
di: Li, Jeffrey, et al.
Pubblicazione: (2025)
di: Li, Jeffrey, et al.
Pubblicazione: (2025)
AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding
di: Chowdhury, Sanjoy, et al.
Pubblicazione: (2025)
di: Chowdhury, Sanjoy, et al.
Pubblicazione: (2025)
Learning from Self Critique and Refinement for Faithful LLM Summarization
di: Hu, Ting-Yao, et al.
Pubblicazione: (2025)
di: Hu, Ting-Yao, et al.
Pubblicazione: (2025)
Dataset Decomposition: Faster LLM Training with Variable Sequence Length Curriculum
di: Pouransari, Hadi, et al.
Pubblicazione: (2024)
di: Pouransari, Hadi, et al.
Pubblicazione: (2024)
FocalLens: Visualizing Narratives through Focalization
di: Alam, S M Raihanul, et al.
Pubblicazione: (2026)
di: Alam, S M Raihanul, et al.
Pubblicazione: (2026)
Graph-Based Captioning: Enhancing Visual Descriptions by Interconnecting Region Captions
di: Hsieh, Yu-Guan, et al.
Pubblicazione: (2024)
di: Hsieh, Yu-Guan, et al.
Pubblicazione: (2024)
Learning to Reason for Hallucination Span Detection
di: Su, Hsuan, et al.
Pubblicazione: (2025)
di: Su, Hsuan, et al.
Pubblicazione: (2025)
Pretraining with hierarchical memories: separating long-tail and common knowledge
di: Pouransari, Hadi, et al.
Pubblicazione: (2025)
di: Pouransari, Hadi, et al.
Pubblicazione: (2025)
Mutual Reinforcement of LLM Dialogue Synthesis and Summarization Capabilities for Few-Shot Dialogue Summarization
di: Lu, Yen-Ju, et al.
Pubblicazione: (2025)
di: Lu, Yen-Ju, et al.
Pubblicazione: (2025)
CatLIP: CLIP-level Visual Recognition Accuracy with 2.7x Faster Pre-training on Web-scale Image-Text Data
di: Mehta, Sachin, et al.
Pubblicazione: (2024)
di: Mehta, Sachin, et al.
Pubblicazione: (2024)
Beyond a Single Extractor: Re-thinking HTML-to-Text Extraction for LLM Pretraining
di: Li, Jeffrey, et al.
Pubblicazione: (2026)
di: Li, Jeffrey, et al.
Pubblicazione: (2026)
Synth4Seg -- Learning Defect Data Synthesis for Defect Segmentation using Bi-level Optimization
di: Mou, Shancong, et al.
Pubblicazione: (2024)
di: Mou, Shancong, et al.
Pubblicazione: (2024)
TrajTok: Learning Trajectory Tokens enables better Video Understanding
di: Zheng, Chenhao, et al.
Pubblicazione: (2026)
di: Zheng, Chenhao, et al.
Pubblicazione: (2026)
SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
di: Brown, Ellis, et al.
Pubblicazione: (2025)
di: Brown, Ellis, et al.
Pubblicazione: (2025)
Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos
di: Seyfioglu, Mehmet Saygin, et al.
Pubblicazione: (2023)
di: Seyfioglu, Mehmet Saygin, et al.
Pubblicazione: (2023)
Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps
di: Chuang, Yung-Sung, et al.
Pubblicazione: (2024)
di: Chuang, Yung-Sung, et al.
Pubblicazione: (2024)
Says Who? Effective Zero-Shot Annotation of Focalization
di: Hicke, Rebecca M. M., et al.
Pubblicazione: (2024)
di: Hicke, Rebecca M. M., et al.
Pubblicazione: (2024)
Deep Exploration of Cross-Lingual Zero-Shot Generalization in Instruction Tuning
di: Han, Janghoon, et al.
Pubblicazione: (2024)
di: Han, Janghoon, et al.
Pubblicazione: (2024)
The Hard Positive Truth about Vision-Language Compositionality
di: Kamath, Amita, et al.
Pubblicazione: (2024)
di: Kamath, Amita, et al.
Pubblicazione: (2024)
EVE: Enabling Anyone to Train Robots using Augmented Reality
di: Wang, Jun, et al.
Pubblicazione: (2024)
di: Wang, Jun, et al.
Pubblicazione: (2024)
LiTo: Surface Light Field Tokenization
di: Chang, Jen-Hao Rick, et al.
Pubblicazione: (2026)
di: Chang, Jen-Hao Rick, et al.
Pubblicazione: (2026)
Learning to Generate Instruction Tuning Datasets for Zero-Shot Task Adaptation
di: Nayak, Nihal V., et al.
Pubblicazione: (2024)
di: Nayak, Nihal V., et al.
Pubblicazione: (2024)
Enabling Natural Zero-Shot Prompting on Encoder Models via Statement-Tuning
di: Elshabrawy, Ahmed, et al.
Pubblicazione: (2024)
di: Elshabrawy, Ahmed, et al.
Pubblicazione: (2024)
Toward Zero-Shot Instruction Following
di: Lou, Renze, et al.
Pubblicazione: (2023)
di: Lou, Renze, et al.
Pubblicazione: (2023)
RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
di: Wu, Yu, et al.
Pubblicazione: (2026)
di: Wu, Yu, et al.
Pubblicazione: (2026)
Assessing LLMs for Zero-shot Abstractive Summarization Through the Lens of Relevance Paraphrasing
di: Askari, Hadi, et al.
Pubblicazione: (2024)
di: Askari, Hadi, et al.
Pubblicazione: (2024)
The Right Time Matters: Data Arrangement Affects Zero-Shot Generalization in Instruction Tuning
di: He, Bingxiang, et al.
Pubblicazione: (2024)
di: He, Bingxiang, et al.
Pubblicazione: (2024)
Velox: Learning Representations of 4D Geometry and Appearance
di: Malik, Anagh, et al.
Pubblicazione: (2026)
di: Malik, Anagh, et al.
Pubblicazione: (2026)
RefDecoder: Enhancing Visual Generation with Conditional Video Decoding
di: Fan, Xiang, et al.
Pubblicazione: (2026)
di: Fan, Xiang, et al.
Pubblicazione: (2026)
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
di: Samragh, Mohammad, et al.
Pubblicazione: (2024)
di: Samragh, Mohammad, et al.
Pubblicazione: (2024)
Documenti analoghi
-
MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
di: Vasu, Pavan Kumar Anasosalu, et al.
Pubblicazione: (2023) -
CLIP with Quality Captions: A Strong Pretraining for Vision Tasks
di: Vasu, Pavan Kumar Anasosalu, et al.
Pubblicazione: (2024) -
SAM-CLIP: Merging Vision Foundation Models towards Semantic and Spatial Understanding
di: Wang, Haoxiang, et al.
Pubblicazione: (2023) -
MobileCLIP2: Improving Multi-Modal Reinforced Training
di: Faghri, Fartash, et al.
Pubblicazione: (2025) -
Knowledge Transfer from Vision Foundation Models for Efficient Training of Small Task-specific Models
di: Vemulapalli, Raviteja, et al.
Pubblicazione: (2023)