VS-LLM: Visual-Semantic Depression Assessment based on LLM for Drawing Projection Test
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Meiqi, Kang, Yaxuan, Li, Xuchen, Hu, Shiyu, Chen, Xiaotang, Kang, Yunfeng, Wang, Weiqiang, Huang, Kaiqi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DTLLM-VLT: Diverse Text Generation for Visual Language Tracking Based on LLM
by: Li, Xuchen, et al.
Published: (2024)
by: Li, Xuchen, et al.
Published: (2024)
DTVLT: A Multi-modal Diverse Text Benchmark for Visual Language Tracking Based on LLM
by: Li, Xuchen, et al.
Published: (2024)
by: Li, Xuchen, et al.
Published: (2024)
Visual Language Tracking with Multi-modal Interaction: A Robust Benchmark
by: Li, Xuchen, et al.
Published: (2024)
by: Li, Xuchen, et al.
Published: (2024)
Finger in Camera Speaks Everything: Unconstrained Air-Writing for Real-World
by: Wu, Meiqi, et al.
Published: (2024)
by: Wu, Meiqi, et al.
Published: (2024)
How Texts Help? A Fine-grained Evaluation to Reveal the Role of Language in Vision-Language Tracking
by: Li, Xuchen, et al.
Published: (2024)
by: Li, Xuchen, et al.
Published: (2024)
Select Less, Reason More: Prioritizing Evidence Purity for Video Reasoning
by: Li, Xuchen, et al.
Published: (2025)
by: Li, Xuchen, et al.
Published: (2025)
CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos
by: Li, Xuchen, et al.
Published: (2025)
by: Li, Xuchen, et al.
Published: (2025)
ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints
by: Wu, Meiqi, et al.
Published: (2025)
by: Wu, Meiqi, et al.
Published: (2025)
MATrack: Efficient Multiscale Adaptive Tracker for Real-Time Nighttime UAV Operations
by: Li, Xuzhao, et al.
Published: (2025)
by: Li, Xuzhao, et al.
Published: (2025)
Beyond Accuracy: Evaluating Grounded Visual Evidence in Thinking with Images
by: Li, Xuchen, et al.
Published: (2026)
by: Li, Xuchen, et al.
Published: (2026)
DARTer: Dynamic Adaptive Representation Tracker for Nighttime UAV Tracking
by: Li, Xuzhao, et al.
Published: (2025)
by: Li, Xuzhao, et al.
Published: (2025)
EduVerse: A User-Defined Multi-Agent Simulation Space for Education Scenario
by: Ma, Yiping, et al.
Published: (2025)
by: Ma, Yiping, et al.
Published: (2025)
FIOVA: A Multi-Annotator Benchmark for Human-Aligned Video Captioning
by: Hu, Shiyu, et al.
Published: (2024)
by: Hu, Shiyu, et al.
Published: (2024)
When LLMs Learn to be Students: The SOEI Framework for Modeling and Evaluating Virtual Student Agents in Educational Interaction
by: Ma, Yiping, et al.
Published: (2024)
by: Ma, Yiping, et al.
Published: (2024)
In Defense of Lazy Visual Grounding for Open-Vocabulary Semantic Segmentation
by: Kang, Dahyun, et al.
Published: (2024)
by: Kang, Dahyun, et al.
Published: (2024)
MagicFuse: Single Image Fusion for Visual and Semantic Reinforcement
by: Zhang, Hao, et al.
Published: (2026)
by: Zhang, Hao, et al.
Published: (2026)
PointT2I: LLM-based text-to-image generation via keypoints
by: Lee, Taekyung, et al.
Published: (2025)
by: Lee, Taekyung, et al.
Published: (2025)
VisualPrompter: Semantic-Aware Prompt Optimization with Visual Feedback for Text-to-Image Synthesis
by: Wu, Shiyu, et al.
Published: (2025)
by: Wu, Shiyu, et al.
Published: (2025)
3DrawAgent: Teaching LLM to Draw in 3D with Early Contrastive Experience
by: Xiao, Hongcan, et al.
Published: (2026)
by: Xiao, Hongcan, et al.
Published: (2026)
Latent Temporal Discrepancy as Motion Prior: A Loss-Weighting Strategy for Dynamic Fidelity in T2V
by: Wu, Meiqi, et al.
Published: (2026)
by: Wu, Meiqi, et al.
Published: (2026)
Assessment of Autism and ADHD: A Comparative Analysis of Drawing Velocity Profiles and the NEPSY Test
by: Fortea-Sevilla, S., et al.
Published: (2024)
by: Fortea-Sevilla, S., et al.
Published: (2024)
SeCG: Semantic-Enhanced 3D Visual Grounding via Cross-modal Graph Attention
by: Xiao, Feng, et al.
Published: (2024)
by: Xiao, Feng, et al.
Published: (2024)
TemCoCo: Temporally Consistent Multi-modal Video Fusion with Visual-Semantic Collaboration
by: Gong, Meiqi, et al.
Published: (2025)
by: Gong, Meiqi, et al.
Published: (2025)
Pick-and-Draw: Training-free Semantic Guidance for Text-to-Image Personalization
by: Lv, Henglei, et al.
Published: (2024)
by: Lv, Henglei, et al.
Published: (2024)
A Deep Learning Framework for Boundary-Aware Semantic Segmentation
by: An, Tai, et al.
Published: (2025)
by: An, Tai, et al.
Published: (2025)
Semantic Visual Simultaneous Localization and Mapping: A Survey
by: Chen, Kaiqi, et al.
Published: (2022)
by: Chen, Kaiqi, et al.
Published: (2022)
Paintings and Drawings Aesthetics Assessment with Rich Attributes for Various Artistic Categories
by: Jin, Xin, et al.
Published: (2024)
by: Jin, Xin, et al.
Published: (2024)
Semantic Draw Engineering for Text-to-Image Creation
by: Li, Yang, et al.
Published: (2023)
by: Li, Yang, et al.
Published: (2023)
MedThink: Explaining Medical Visual Question Answering via Multimodal Decision-Making Rationale
by: Gai, Xiaotang, et al.
Published: (2024)
by: Gai, Xiaotang, et al.
Published: (2024)
Cross-Stage Attention Propagation for Efficient Semantic Segmentation
by: Kang, Beoungwoo
Published: (2026)
by: Kang, Beoungwoo
Published: (2026)
Towards Semantic Equivalence of Tokenization in Multimodal LLM
by: Wu, Shengqiong, et al.
Published: (2024)
by: Wu, Shengqiong, et al.
Published: (2024)
DiffVP: Differential Visual Semantic Prompting for LLM-Based CT Report Generation
by: Tian, Yuhe, et al.
Published: (2026)
by: Tian, Yuhe, et al.
Published: (2026)
Does VLM Classification Benefit from LLM Description Semantics?
by: Ma, Pingchuan, et al.
Published: (2024)
by: Ma, Pingchuan, et al.
Published: (2024)
ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
by: Hu, Xiwei, et al.
Published: (2024)
by: Hu, Xiwei, et al.
Published: (2024)
KeyframeFace: Language-Driven Facial Animation via Semantic Keyframes
by: Wu, Jingchao, et al.
Published: (2025)
by: Wu, Jingchao, et al.
Published: (2025)
LocLLM: Exploiting Generalizable Human Keypoint Localization via Large Language Model
by: Wang, Dongkai, et al.
Published: (2024)
by: Wang, Dongkai, et al.
Published: (2024)
WIPES: Wavelet-based Visual Primitives
by: Zhang, Wenhao, et al.
Published: (2025)
by: Zhang, Wenhao, et al.
Published: (2025)
MergeSAM: Unsupervised change detection of remote sensing images based on the Segment Anything Model
by: Hu, Meiqi, et al.
Published: (2025)
by: Hu, Meiqi, et al.
Published: (2025)
SemanticFace: Semantic Facial Action Estimation via Semantic Distillation in Interpretable Space
by: Kang, Zejian, et al.
Published: (2026)
by: Kang, Zejian, et al.
Published: (2026)
Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning
by: Li, Xuchen, et al.
Published: (2025)
by: Li, Xuchen, et al.
Published: (2025)
Similar Items
-
DTLLM-VLT: Diverse Text Generation for Visual Language Tracking Based on LLM
by: Li, Xuchen, et al.
Published: (2024) -
DTVLT: A Multi-modal Diverse Text Benchmark for Visual Language Tracking Based on LLM
by: Li, Xuchen, et al.
Published: (2024) -
Visual Language Tracking with Multi-modal Interaction: A Robust Benchmark
by: Li, Xuchen, et al.
Published: (2024) -
Finger in Camera Speaks Everything: Unconstrained Air-Writing for Real-World
by: Wu, Meiqi, et al.
Published: (2024) -
How Texts Help? A Fine-grained Evaluation to Reveal the Role of Language in Vision-Language Tracking
by: Li, Xuchen, et al.
Published: (2024)