Improving Brain-to-Image Reconstruction via Fine-Grained Text Bridging
Fuente:
arXiv
Salvato in:
| Autori principali: | Xia, Runze, Feng, Shuo, Wang, Renzhi, Yin, Congchi, Wen, Xuyun, Li, Piji |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Improve Language Model and Brain Alignment via Associative Memory
di: Yin, Congchi, et al.
Pubblicazione: (2025)
di: Yin, Congchi, et al.
Pubblicazione: (2025)
Decoding the Echoes of Vision from fMRI: Memory Disentangling for Past Semantic Information
di: Xia, Runze, et al.
Pubblicazione: (2024)
di: Xia, Runze, et al.
Pubblicazione: (2024)
Language Reconstruction with Brain Predictive Coding from fMRI Data
di: Yin, Congchi, et al.
Pubblicazione: (2024)
di: Yin, Congchi, et al.
Pubblicazione: (2024)
Medical Image Synthesis via Fine-Grained Image-Text Alignment and Anatomy-Pathology Prompting
di: Chen, Wenting, et al.
Pubblicazione: (2024)
di: Chen, Wenting, et al.
Pubblicazione: (2024)
Rethinking Cross-Subject Data Splitting for Brain-to-Text Decoding
di: Yin, Congchi, et al.
Pubblicazione: (2023)
di: Yin, Congchi, et al.
Pubblicazione: (2023)
CompAlign: Improving Compositional Text-to-Image Generation with a Complex Benchmark and Fine-Grained Feedback
di: Wan, Yixin, et al.
Pubblicazione: (2025)
di: Wan, Yixin, et al.
Pubblicazione: (2025)
Fine-Grained Image-Text Alignment in Medical Imaging Enables Explainable Cyclic Image-Report Generation
di: Chen, Wenting, et al.
Pubblicazione: (2023)
di: Chen, Wenting, et al.
Pubblicazione: (2023)
Enhancing Fine-Grained Image Classifications via Cascaded Vision Language Models
di: Wei, Canshi
Pubblicazione: (2024)
di: Wei, Canshi
Pubblicazione: (2024)
HELPD: Mitigating Hallucination of LVLMs by Hierarchical Feedback Learning with Vision-enhanced Penalty Decoding
di: Yuan, Fan, et al.
Pubblicazione: (2024)
di: Yuan, Fan, et al.
Pubblicazione: (2024)
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
di: Shen, Yifan, et al.
Pubblicazione: (2025)
di: Shen, Yifan, et al.
Pubblicazione: (2025)
Improving Text-to-Image Consistency via Automatic Prompt Optimization
di: Mañas, Oscar, et al.
Pubblicazione: (2024)
di: Mañas, Oscar, et al.
Pubblicazione: (2024)
TROPE: TRaining-Free Object-Part Enhancement for Seamlessly Improving Fine-Grained Zero-Shot Image Captioning
di: Feinglass, Joshua, et al.
Pubblicazione: (2024)
di: Feinglass, Joshua, et al.
Pubblicazione: (2024)
FineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text Inputs
di: Asokan, Mothilal, et al.
Pubblicazione: (2025)
di: Asokan, Mothilal, et al.
Pubblicazione: (2025)
Beyond Thumbs Up/Down: Untangling Challenges of Fine-Grained Feedback for Text-to-Image Generation
di: Collins, Katherine M., et al.
Pubblicazione: (2024)
di: Collins, Katherine M., et al.
Pubblicazione: (2024)
VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
di: Liu, Peng, et al.
Pubblicazione: (2025)
di: Liu, Peng, et al.
Pubblicazione: (2025)
LLM-based Hierarchical Concept Decomposition for Interpretable Fine-Grained Image Classification
di: Qu, Renyi, et al.
Pubblicazione: (2024)
di: Qu, Renyi, et al.
Pubblicazione: (2024)
Multi-Modal 3D Mesh Reconstruction from Images and Text
di: Reka, Melvin, et al.
Pubblicazione: (2025)
di: Reka, Melvin, et al.
Pubblicazione: (2025)
Text-Video Retrieval via Variational Multi-Modal Hypergraph Networks
di: Li, Qian, et al.
Pubblicazione: (2024)
di: Li, Qian, et al.
Pubblicazione: (2024)
Human-Aligned MLLM Judges for Fine-Grained Image Editing Evaluation: A Benchmark, Framework, and Analysis
di: Liu, Runzhou, et al.
Pubblicazione: (2026)
di: Liu, Runzhou, et al.
Pubblicazione: (2026)
Semantic are Beacons: A Semantic Perspective for Unveiling Parameter-Efficient Fine-Tuning in Knowledge Learning
di: Wang, Renzhi, et al.
Pubblicazione: (2024)
di: Wang, Renzhi, et al.
Pubblicazione: (2024)
FINEMATCH: Aspect-based Fine-grained Image and Text Mismatch Detection and Correction
di: Hua, Hang, et al.
Pubblicazione: (2024)
di: Hua, Hang, et al.
Pubblicazione: (2024)
MVAM: Multi-View Attention Method for Fine-grained Image-Text Matching
di: Cui, Wanqing, et al.
Pubblicazione: (2024)
di: Cui, Wanqing, et al.
Pubblicazione: (2024)
EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models
di: Tan, Zhiyu, et al.
Pubblicazione: (2024)
di: Tan, Zhiyu, et al.
Pubblicazione: (2024)
NAVIG: Natural Language-guided Analysis with Vision Language Models for Image Geo-localization
di: Zhang, Zheyuan, et al.
Pubblicazione: (2025)
di: Zhang, Zheyuan, et al.
Pubblicazione: (2025)
MedRCube: A Multidimensional Framework for Fine-Grained and In-Depth Evaluation of MLLMs in Medical Imaging
di: Bao, Zhijie, et al.
Pubblicazione: (2026)
di: Bao, Zhijie, et al.
Pubblicazione: (2026)
FineCops-Ref: A new Dataset and Task for Fine-Grained Compositional Referring Expression Comprehension
di: Liu, Junzhuo, et al.
Pubblicazione: (2024)
di: Liu, Junzhuo, et al.
Pubblicazione: (2024)
Generative Visual Commonsense Answering and Explaining with Generative Scene Graph Constructing
di: Yuan, Fan, et al.
Pubblicazione: (2025)
di: Yuan, Fan, et al.
Pubblicazione: (2025)
MC-MKE: A Fine-Grained Multimodal Knowledge Editing Benchmark Emphasizing Modality Consistency
di: Zhang, Junzhe, et al.
Pubblicazione: (2024)
di: Zhang, Junzhe, et al.
Pubblicazione: (2024)
VisRet: Visualization Improves Knowledge-Intensive Text-to-Image Retrieval
di: Wu, Di, et al.
Pubblicazione: (2025)
di: Wu, Di, et al.
Pubblicazione: (2025)
RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction
di: Wang, Yuchi, et al.
Pubblicazione: (2025)
di: Wang, Yuchi, et al.
Pubblicazione: (2025)
Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Models
di: Yamabe, Shojiro, et al.
Pubblicazione: (2025)
di: Yamabe, Shojiro, et al.
Pubblicazione: (2025)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
di: Wu, Yin, et al.
Pubblicazione: (2025)
di: Wu, Yin, et al.
Pubblicazione: (2025)
You May Speak Freely: Improving the Fine-Grained Visual Recognition Capabilities of Multimodal Large Language Models with Answer Extraction
di: Lawrence, Logan, et al.
Pubblicazione: (2025)
di: Lawrence, Logan, et al.
Pubblicazione: (2025)
Fine-Grained Evaluation of Large Vision-Language Models in Autonomous Driving
di: Li, Yue, et al.
Pubblicazione: (2025)
di: Li, Yue, et al.
Pubblicazione: (2025)
Rethinking Composed Image Retrieval Evaluation: A Fine-Grained Benchmark from Image Editing
di: Song, Tingyu, et al.
Pubblicazione: (2026)
di: Song, Tingyu, et al.
Pubblicazione: (2026)
Learning from Fine-Grained Visual Discrepancies: Mitigating Multimodal Hallucinations via In-Context Visual Contrastive Optimization
di: Deng, Haolin, et al.
Pubblicazione: (2026)
di: Deng, Haolin, et al.
Pubblicazione: (2026)
Praxis-VLM: Vision-Grounded Decision Making via Text-Driven Reinforcement Learning
di: Hu, Zhe, et al.
Pubblicazione: (2025)
di: Hu, Zhe, et al.
Pubblicazione: (2025)
Optimizing Prompts for Text-to-Image Generation
di: Hao, Yaru, et al.
Pubblicazione: (2022)
di: Hao, Yaru, et al.
Pubblicazione: (2022)
Fine-grained Spatiotemporal Grounding on Egocentric Videos
di: Liang, Shuo, et al.
Pubblicazione: (2025)
di: Liang, Shuo, et al.
Pubblicazione: (2025)
STEMTOX: From Social Tags to Fine-Grained Toxic Meme Detection via Entropy-Guided Multi-Task Learning
di: Swain, Subhankar, et al.
Pubblicazione: (2025)
di: Swain, Subhankar, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Improve Language Model and Brain Alignment via Associative Memory
di: Yin, Congchi, et al.
Pubblicazione: (2025) -
Decoding the Echoes of Vision from fMRI: Memory Disentangling for Past Semantic Information
di: Xia, Runze, et al.
Pubblicazione: (2024) -
Language Reconstruction with Brain Predictive Coding from fMRI Data
di: Yin, Congchi, et al.
Pubblicazione: (2024) -
Medical Image Synthesis via Fine-Grained Image-Text Alignment and Anatomy-Pathology Prompting
di: Chen, Wenting, et al.
Pubblicazione: (2024) -
Rethinking Cross-Subject Data Splitting for Brain-to-Text Decoding
di: Yin, Congchi, et al.
Pubblicazione: (2023)