SliceWorld: A Predictive and Controllable World-State Model for CT Report Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Tian, Yuanhe, Song, Yan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation
by: Tian, Yuanhe, et al.
Published: (2025)
by: Tian, Yuanhe, et al.
Published: (2025)
A Systematic Review of Deep Learning-based Research on Radiology Report Generation
by: Liu, Chang, et al.
Published: (2023)
by: Liu, Chang, et al.
Published: (2023)
Computed Tomography Visual Question Answering with Cross-modal Feature Graphing
by: Tian, Yuanhe, et al.
Published: (2025)
by: Tian, Yuanhe, et al.
Published: (2025)
Open World Scene Graph Generation using Vision Language Models
by: Dutta, Amartya, et al.
Published: (2025)
by: Dutta, Amartya, et al.
Published: (2025)
Pandora: Towards General World Model with Natural Language Actions and Video States
by: Xiang, Jiannan, et al.
Published: (2024)
by: Xiang, Jiannan, et al.
Published: (2024)
Exploring In-Image Machine Translation with Real-World Background
by: Tian, Yanzhi, et al.
Published: (2025)
by: Tian, Yanzhi, et al.
Published: (2025)
From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction
by: Zhao, Zhida, et al.
Published: (2025)
by: Zhao, Zhida, et al.
Published: (2025)
SliceFine: The Universal Winning-Slice Hypothesis for Pretrained Networks
by: Kowsher, Md, et al.
Published: (2025)
by: Kowsher, Md, et al.
Published: (2025)
CTest-Metric: A Unified Framework to Assess Clinical Validity of Metrics for CT Report Generation
by: Sharma, Vanshali, et al.
Published: (2026)
by: Sharma, Vanshali, et al.
Published: (2026)
"See the World, Discover Knowledge": A Chinese Factuality Evaluation for Large Vision Language Models
by: Gu, Jihao, et al.
Published: (2025)
by: Gu, Jihao, et al.
Published: (2025)
EyeWorld: A Generative World Model of Ocular State and Dynamics
by: Gao, Ziyu, et al.
Published: (2026)
by: Gao, Ziyu, et al.
Published: (2026)
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
by: Liu, Chengzhi, et al.
Published: (2026)
by: Liu, Chengzhi, et al.
Published: (2026)
How Far Are Vision-Language Models from Constructing the Real World? A Benchmark for Physical Generative Reasoning
by: Yang, Luyu, et al.
Published: (2026)
by: Yang, Luyu, et al.
Published: (2026)
Web World Models
by: Feng, Jichen, et al.
Published: (2025)
by: Feng, Jichen, et al.
Published: (2025)
PAN: A World Model for General, Interactable, and Long-Horizon World Simulation
by: PAN Team, et al.
Published: (2025)
by: PAN Team, et al.
Published: (2025)
Code2World: A GUI World Model via Renderable Code Generation
by: Zheng, Yuhao, et al.
Published: (2026)
by: Zheng, Yuhao, et al.
Published: (2026)
Multi-Dimensional Insights: Benchmarking Real-World Personalization in Large Multimodal Models
by: Zhang, YiFan, et al.
Published: (2024)
by: Zhang, YiFan, et al.
Published: (2024)
What-If World: A Causal Benchmark for General World Models in Embodied Scenarios
by: Cai, Kunlin, et al.
Published: (2026)
by: Cai, Kunlin, et al.
Published: (2026)
Generating Reports or Repeating Templates? Measuring and Mitigating Template Collapse in 3D CT Report Generation
by: Maye-Lasserre, Tom, et al.
Published: (2026)
by: Maye-Lasserre, Tom, et al.
Published: (2026)
Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model
by: Huang, Haoyang, et al.
Published: (2025)
by: Huang, Haoyang, et al.
Published: (2025)
3D-VLA: A 3D Vision-Language-Action Generative World Model
by: Zhen, Haoyu, et al.
Published: (2024)
by: Zhen, Haoyu, et al.
Published: (2024)
World Action Models: The Next Frontier in Embodied AI
by: Wang, Siyin, et al.
Published: (2026)
by: Wang, Siyin, et al.
Published: (2026)
DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models
by: Kim, Sungnyun, et al.
Published: (2024)
by: Kim, Sungnyun, et al.
Published: (2024)
MathScape: Benchmarking Multimodal Large Language Models in Real-World Mathematical Contexts
by: Liang, Hao, et al.
Published: (2024)
by: Liang, Hao, et al.
Published: (2024)
Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models
by: Bitton-Guetta, Nitzan, et al.
Published: (2024)
by: Bitton-Guetta, Nitzan, et al.
Published: (2024)
AI Sees Your Location, But With A Bias Toward The Wealthy World
by: Huang, Jingyuan, et al.
Published: (2025)
by: Huang, Jingyuan, et al.
Published: (2025)
CT-GLIP: 3D Grounded Language-Image Pretraining with CT Scans and Radiology Reports for Full-Body Scenarios
by: Lin, Jingyang, et al.
Published: (2024)
by: Lin, Jingyang, et al.
Published: (2024)
LaB-RAG: Label Boosted Retrieval Augmented Generation for Radiology Report Generation
by: Song, Steven, et al.
Published: (2024)
by: Song, Steven, et al.
Published: (2024)
Rethinking Radiology Report Generation via Causal Inspired Counterfactual Augmentation
by: Song, Xiao, et al.
Published: (2023)
by: Song, Xiao, et al.
Published: (2023)
LADDER: Language-Driven Slice Discovery and Error Rectification in Vision Classifiers
by: Ghosh, Shantanu, et al.
Published: (2024)
by: Ghosh, Shantanu, et al.
Published: (2024)
Can World Models Benefit VLMs for World Dynamics?
by: Zhang, Kevin, et al.
Published: (2025)
by: Zhang, Kevin, et al.
Published: (2025)
PixelWorld: How Far Are We from Perceiving Everything as Pixels?
by: Lyu, Zhiheng, et al.
Published: (2025)
by: Lyu, Zhiheng, et al.
Published: (2025)
ICON: Improving Inter-Report Consistency in Radiology Report Generation via Lesion-aware Mixup Augmentation
by: Hou, Wenjun, et al.
Published: (2024)
by: Hou, Wenjun, et al.
Published: (2024)
RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation
by: Li, Huiqiong, et al.
Published: (2026)
by: Li, Huiqiong, et al.
Published: (2026)
WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens
by: Wang, Xiaofeng, et al.
Published: (2024)
by: Wang, Xiaofeng, et al.
Published: (2024)
We Should Chart an Atlas of All the World's Models
by: Horwitz, Eliahu, et al.
Published: (2025)
by: Horwitz, Eliahu, et al.
Published: (2025)
PhoStream: Benchmarking Real-World Streaming for Omnimodal Assistants in Mobile Scenarios
by: Lu, Xudong, et al.
Published: (2026)
by: Lu, Xudong, et al.
Published: (2026)
ChartCheck: Explainable Fact-Checking over Real-World Chart Images
by: Akhtar, Mubashara, et al.
Published: (2023)
by: Akhtar, Mubashara, et al.
Published: (2023)
RADAR: Enhancing Radiology Report Generation with Supplementary Knowledge Injection
by: Hou, Wenjun, et al.
Published: (2025)
by: Hou, Wenjun, et al.
Published: (2025)
Privacy-Aware Camera 2.0 Technical Report
by: Song, Huan, et al.
Published: (2026)
by: Song, Huan, et al.
Published: (2026)
Similar Items
-
Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation
by: Tian, Yuanhe, et al.
Published: (2025) -
A Systematic Review of Deep Learning-based Research on Radiology Report Generation
by: Liu, Chang, et al.
Published: (2023) -
Computed Tomography Visual Question Answering with Cross-modal Feature Graphing
by: Tian, Yuanhe, et al.
Published: (2025) -
Open World Scene Graph Generation using Vision Language Models
by: Dutta, Amartya, et al.
Published: (2025) -
Pandora: Towards General World Model with Natural Language Actions and Video States
by: Xiang, Jiannan, et al.
Published: (2024)