Learning Sparse Visual Representations via Spatial-Semantic Factorization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Theodore Zhengde, Kiblawi, Sid, Yang, Jianwei, Usuyama, Naoto, Tan, Reuben, Codella, Noel C, Naumann, Tristan, Poon, Hoifung, Wei, Mu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Boltzmann Attention Sampling for Image Analysis with Small Objects
von: Zhao, Theodore, et al.
Veröffentlicht: (2025)
von: Zhao, Theodore, et al.
Veröffentlicht: (2025)
Exploring Scaling Laws for EHR Foundation Models
von: Zhang, Sheng, et al.
Veröffentlicht: (2025)
von: Zhang, Sheng, et al.
Veröffentlicht: (2025)
X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains
von: Liu, Qianchu, et al.
Veröffentlicht: (2025)
von: Liu, Qianchu, et al.
Veröffentlicht: (2025)
BiomedParse: a biomedical foundation model for image parsing of everything everywhere all at once
von: Zhao, Theodore, et al.
Veröffentlicht: (2024)
von: Zhao, Theodore, et al.
Veröffentlicht: (2024)
Med-RLVR: Emerging Medical Reasoning from a 3B base model via reinforcement Learning
von: Zhang, Sheng, et al.
Veröffentlicht: (2025)
von: Zhang, Sheng, et al.
Veröffentlicht: (2025)
Pareto Optimal Learning for Estimating Large Language Model Errors
von: Zhao, Theodore, et al.
Veröffentlicht: (2023)
von: Zhao, Theodore, et al.
Veröffentlicht: (2023)
Scaling medical imaging report generation with multimodal reinforcement learning
von: Liu, Qianchu, et al.
Veröffentlicht: (2026)
von: Liu, Qianchu, et al.
Veröffentlicht: (2026)
OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning
von: Ossowski, Timothy, et al.
Veröffentlicht: (2025)
von: Ossowski, Timothy, et al.
Veröffentlicht: (2025)
Foundation Models for Biomedical Image Segmentation: A Survey
von: Lee, Ho Hin, et al.
Veröffentlicht: (2024)
von: Lee, Ho Hin, et al.
Veröffentlicht: (2024)
Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration
von: Huang, James Y., et al.
Veröffentlicht: (2025)
von: Huang, James Y., et al.
Veröffentlicht: (2025)
Semantic-Clipping: Efficient Vision-Language Modeling with Semantic-Guidedd Visual Selection
von: Li, Bangzheng, et al.
Veröffentlicht: (2025)
von: Li, Bangzheng, et al.
Veröffentlicht: (2025)
Attribute Structuring Improves LLM-Based Evaluation of Clinical Text Summaries
von: Gero, Zelalem, et al.
Veröffentlicht: (2024)
von: Gero, Zelalem, et al.
Veröffentlicht: (2024)
DocLens: Multi-aspect Fine-grained Evaluation for Medical Text Generation
von: Xie, Yiqing, et al.
Veröffentlicht: (2023)
von: Xie, Yiqing, et al.
Veröffentlicht: (2023)
Universal Abstraction: Harnessing Frontier Models to Structure Real-World Data at Scale
von: Wong, Cliff, et al.
Veröffentlicht: (2025)
von: Wong, Cliff, et al.
Veröffentlicht: (2025)
AURAD: Anatomy-Pathology Unified Radiology Synthesis with Progressive Representations
von: Ding, Shuhan, et al.
Veröffentlicht: (2025)
von: Ding, Shuhan, et al.
Veröffentlicht: (2025)
The Illusion of Readiness in Health AI
von: Gu, Yu, et al.
Veröffentlicht: (2025)
von: Gu, Yu, et al.
Veröffentlicht: (2025)
Generative Medical Event Models Improve with Scale
von: Waxler, Shane, et al.
Veröffentlicht: (2025)
von: Waxler, Shane, et al.
Veröffentlicht: (2025)
BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs
von: Zhang, Sheng, et al.
Veröffentlicht: (2023)
von: Zhang, Sheng, et al.
Veröffentlicht: (2023)
CancerGUIDE: Cancer Guideline Understanding via Internal Disagreement Estimation
von: Unell, Alyssa, et al.
Veröffentlicht: (2025)
von: Unell, Alyssa, et al.
Veröffentlicht: (2025)
Cautionary Tales on Synthetic Controls in Survival Analyses
von: Curth, Alicia, et al.
Veröffentlicht: (2023)
von: Curth, Alicia, et al.
Veröffentlicht: (2023)
Enhancing Visual Representation with Textual Semantics: Textual Semantics-Powered Prototypes for Heterogeneous Federated Learning
von: Wu, Xinghao, et al.
Veröffentlicht: (2025)
von: Wu, Xinghao, et al.
Veröffentlicht: (2025)
Towards a clinically accessible radiology foundation model: open-access and lightweight, with automated evaluation
von: Chaves, Juan Manuel Zambrano, et al.
Veröffentlicht: (2024)
von: Chaves, Juan Manuel Zambrano, et al.
Veröffentlicht: (2024)
UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity Recognition
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2023)
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2023)
From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning
von: Xu, Nan, et al.
Veröffentlicht: (2024)
von: Xu, Nan, et al.
Veröffentlicht: (2024)
ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
von: Liu, Qin, et al.
Veröffentlicht: (2025)
von: Liu, Qin, et al.
Veröffentlicht: (2025)
Robust Hyperspectral Image Panshapring via Sparse Spatial-Spectral Representation
von: Lee, Chia-Ming, et al.
Veröffentlicht: (2025)
von: Lee, Chia-Ming, et al.
Veröffentlicht: (2025)
Fully Authentic Visual Question Answering Dataset from Online Communities
von: Chen, Chongyan, et al.
Veröffentlicht: (2023)
von: Chen, Chongyan, et al.
Veröffentlicht: (2023)
Energy Management of V2G‐Containing Multiource Microgrid Cluster Based on Two‐Layer Hybrid Game
von: Mei Li, et al.
Veröffentlicht: (2025)
von: Mei Li, et al.
Veröffentlicht: (2025)
MindJourney: Test-Time Scaling with World Models for Spatial Reasoning
von: Yang, Yuncong, et al.
Veröffentlicht: (2025)
von: Yang, Yuncong, et al.
Veröffentlicht: (2025)
NeRF-VO: Real-Time Sparse Visual Odometry with Neural Radiance Fields
von: Naumann, Jens, et al.
Veröffentlicht: (2023)
von: Naumann, Jens, et al.
Veröffentlicht: (2023)
SITE: towards Spatial Intelligence Thorough Evaluation
von: Wang, Wenqi, et al.
Veröffentlicht: (2025)
von: Wang, Wenqi, et al.
Veröffentlicht: (2025)
Neural Image Compression Using Masked Sparse Visual Representation
von: Jiang, Wei, et al.
Veröffentlicht: (2023)
von: Jiang, Wei, et al.
Veröffentlicht: (2023)
A Theoretical Framework for Visual Weight: Contrast, Size, and Edge Sharpness as Candidate Determinants of Perceived Dominance in Static Images
von: Vasandani, Sid
Veröffentlicht: (2026)
von: Vasandani, Sid
Veröffentlicht: (2026)
TRIALSCOPE: A Unifying Causal Framework for Scaling Real-World Evidence Generation with Biomedical Language Models
von: González, Javier, et al.
Veröffentlicht: (2023)
von: González, Javier, et al.
Veröffentlicht: (2023)
Diversity Over Frequency: Rethinking Tool Use in Visual Chain-of-Thought Agents
von: Kim, Dong-Hee, et al.
Veröffentlicht: (2026)
von: Kim, Dong-Hee, et al.
Veröffentlicht: (2026)
Spin-Transfer-Torque Induced Spatially Nonuniform Switching in Ferrimagnets
von: Zhang, Xue, et al.
Veröffentlicht: (2024)
von: Zhang, Xue, et al.
Veröffentlicht: (2024)
Visualizing Spatial Semantics of Dimensionally Reduced Text Embeddings
von: Liu, Wei, et al.
Veröffentlicht: (2024)
von: Liu, Wei, et al.
Veröffentlicht: (2024)
Naked Clams: A Comprehensive Analysis of Their Global Potential for Commercial Aquaculture
von: Jia Rong Poon, et al.
Veröffentlicht: (2025)
von: Jia Rong Poon, et al.
Veröffentlicht: (2025)
Exponential Localization of Spatial Random Permutations in One Dimension
von: Drogin, Reuben, et al.
Veröffentlicht: (2026)
von: Drogin, Reuben, et al.
Veröffentlicht: (2026)
Revisiting the Analytical Solution of Spin-Orbit Torque Switched Nanoscale Perpendicular Ferromagnet
von: Zhang, Xue, et al.
Veröffentlicht: (2024)
von: Zhang, Xue, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Boltzmann Attention Sampling for Image Analysis with Small Objects
von: Zhao, Theodore, et al.
Veröffentlicht: (2025) -
Exploring Scaling Laws for EHR Foundation Models
von: Zhang, Sheng, et al.
Veröffentlicht: (2025) -
X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains
von: Liu, Qianchu, et al.
Veröffentlicht: (2025) -
BiomedParse: a biomedical foundation model for image parsing of everything everywhere all at once
von: Zhao, Theodore, et al.
Veröffentlicht: (2024) -
Med-RLVR: Emerging Medical Reasoning from a 3B base model via reinforcement Learning
von: Zhang, Sheng, et al.
Veröffentlicht: (2025)