Integrating Large Language Models into a Tri-Modal Architecture for Automated Depression Classification on the DAIC-WOZ
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Patapati, Santosh V. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PhysNav-DG: A Novel Adaptive Framework for Robust VLM-Sensor Fusion in Navigation Applications
von: Srinivasan, Trisanth, et al.
Veröffentlicht: (2025)
von: Srinivasan, Trisanth, et al.
Veröffentlicht: (2025)
RMAdapter: Reconstruction-based Multi-Modal Adapter for Vision-Language Models
von: Lin, Xiang, et al.
Veröffentlicht: (2025)
von: Lin, Xiang, et al.
Veröffentlicht: (2025)
Goal-Based Vision-Language Driving
von: Patapati, Santosh, et al.
Veröffentlicht: (2025)
von: Patapati, Santosh, et al.
Veröffentlicht: (2025)
PlanLLM: Video Procedure Planning with Refinable Large Language Models
von: Yang, Dejie, et al.
Veröffentlicht: (2024)
von: Yang, Dejie, et al.
Veröffentlicht: (2024)
Towards Multi-Task Multi-Modal Models: A Video Generative Perspective
von: Yu, Lijun
Veröffentlicht: (2024)
von: Yu, Lijun
Veröffentlicht: (2024)
Vision-Language Meets the Skeleton: Progressively Distillation with Cross-Modal Knowledge for 3D Action Representation Learning
von: Chen, Yang, et al.
Veröffentlicht: (2024)
von: Chen, Yang, et al.
Veröffentlicht: (2024)
Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language Models
von: Zhou, Shengli, et al.
Veröffentlicht: (2026)
von: Zhou, Shengli, et al.
Veröffentlicht: (2026)
Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis
von: Huang, Po-Hsuan, et al.
Veröffentlicht: (2024)
von: Huang, Po-Hsuan, et al.
Veröffentlicht: (2024)
ReconBoost: Boosting Can Achieve Modality Reconcilement
von: Hua, Cong, et al.
Veröffentlicht: (2024)
von: Hua, Cong, et al.
Veröffentlicht: (2024)
CLARGA: Multimodal Graph Representation Learning over Arbitrary Sets of Modalities
von: Patapati, Santosh
Veröffentlicht: (2025)
von: Patapati, Santosh
Veröffentlicht: (2025)
TbExplain: A Text-based Explanation Method for Scene Classification Models with the Statistical Prediction Correction
von: Aminimehr, Amirhossein, et al.
Veröffentlicht: (2023)
von: Aminimehr, Amirhossein, et al.
Veröffentlicht: (2023)
OneLLM: One Framework to Align All Modalities with Language
von: Han, Jiaming, et al.
Veröffentlicht: (2023)
von: Han, Jiaming, et al.
Veröffentlicht: (2023)
COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity Recognition
von: Chen, Baiyu, et al.
Veröffentlicht: (2025)
von: Chen, Baiyu, et al.
Veröffentlicht: (2025)
Can We Edit Multimodal Large Language Models?
von: Cheng, Siyuan, et al.
Veröffentlicht: (2023)
von: Cheng, Siyuan, et al.
Veröffentlicht: (2023)
HER2 Expression Prediction with Flexible Multi-Modal Inputs via Dynamic Bidirectional Reconstruction
von: Qin, Jie, et al.
Veröffentlicht: (2025)
von: Qin, Jie, et al.
Veröffentlicht: (2025)
An Efficient and Explanatory Image and Text Clustering System with Multimodal Autoencoder Architecture
von: Shi, Tiancheng, et al.
Veröffentlicht: (2024)
von: Shi, Tiancheng, et al.
Veröffentlicht: (2024)
Squeezing Capacity from Multimodal Large Language Models for Subject-driven Generation
von: Zheng, Shuhong, et al.
Veröffentlicht: (2026)
von: Zheng, Shuhong, et al.
Veröffentlicht: (2026)
HeCoFuse: Cross-Modal Complementary V2X Cooperative Perception with Heterogeneous Sensors
von: Wei, Chuheng, et al.
Veröffentlicht: (2025)
von: Wei, Chuheng, et al.
Veröffentlicht: (2025)
Understanding the Fine-Grained Knowledge Capabilities of Vision-Language Models
von: Ghosh, Dhruba, et al.
Veröffentlicht: (2026)
von: Ghosh, Dhruba, et al.
Veröffentlicht: (2026)
To Forget or Not? Towards Practical Knowledge Unlearning for Large Language Models
von: Tian, Bozhong, et al.
Veröffentlicht: (2024)
von: Tian, Bozhong, et al.
Veröffentlicht: (2024)
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models
von: Chen, Wenting, et al.
Veröffentlicht: (2025)
von: Chen, Wenting, et al.
Veröffentlicht: (2025)
IoT-LM: Large Multisensory Language Models for the Internet of Things
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
RAP: Retrieval-Augmented Personalization for Multimodal Large Language Models
von: Hao, Haoran, et al.
Veröffentlicht: (2024)
von: Hao, Haoran, et al.
Veröffentlicht: (2024)
Reducing Hallucinations in Vision-Language Models via Latent Space Steering
von: Liu, Sheng, et al.
Veröffentlicht: (2024)
von: Liu, Sheng, et al.
Veröffentlicht: (2024)
Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models
von: Zhang, Yabin, et al.
Veröffentlicht: (2024)
von: Zhang, Yabin, et al.
Veröffentlicht: (2024)
David and Goliath: Small One-step Model Beats Large Diffusion with Score Post-training
von: Luo, Weijian, et al.
Veröffentlicht: (2024)
von: Luo, Weijian, et al.
Veröffentlicht: (2024)
PAND: Prompt-Aware Neighborhood Distillation for Lightweight Fine-Grained Visual Classification
von: Luo, Qiuming, et al.
Veröffentlicht: (2026)
von: Luo, Qiuming, et al.
Veröffentlicht: (2026)
Deciphering Personalization: Towards Fine-Grained Explainability in Natural Language for Personalized Image Generation Models
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes?
von: Yao, Yang, et al.
Veröffentlicht: (2025)
von: Yao, Yang, et al.
Veröffentlicht: (2025)
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models
von: Sun, Zeyi, et al.
Veröffentlicht: (2024)
von: Sun, Zeyi, et al.
Veröffentlicht: (2024)
From Linguistic Giants to Sensory Maestros: A Survey on Cross-Modal Reasoning with Large Language Models
von: Qian, Shengsheng, et al.
Veröffentlicht: (2024)
von: Qian, Shengsheng, et al.
Veröffentlicht: (2024)
PC$^2$: Pseudo-Classification Based Pseudo-Captioning for Noisy Correspondence Learning in Cross-Modal Retrieval
von: Duan, Yue, et al.
Veröffentlicht: (2024)
von: Duan, Yue, et al.
Veröffentlicht: (2024)
CLIP-MG: Guiding Semantic Attention with Skeletal Pose Features and RGB Data for Micro-Gesture Recognition on the iMiGUE Dataset
von: Patapati, Santosh, et al.
Veröffentlicht: (2025)
von: Patapati, Santosh, et al.
Veröffentlicht: (2025)
Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding
von: Deng, Ailin, et al.
Veröffentlicht: (2024)
von: Deng, Ailin, et al.
Veröffentlicht: (2024)
COSMIC: Clique-Oriented Semantic Multi-space Integration for Robust CLIP Test-Time Adaptation
von: Huang, Fanding, et al.
Veröffentlicht: (2025)
von: Huang, Fanding, et al.
Veröffentlicht: (2025)
Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models
von: Yu, Xiaomin, et al.
Veröffentlicht: (2026)
von: Yu, Xiaomin, et al.
Veröffentlicht: (2026)
OmniEvalKit: A Modular, Lightweight Toolbox for Evaluating Large Language Model and its Omni-Extensions
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2024)
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2024)
Continuous Sign Language Recognition System using Deep Learning with MediaPipe Holistic
von: Srivastava, Sharvani, et al.
Veröffentlicht: (2024)
von: Srivastava, Sharvani, et al.
Veröffentlicht: (2024)
Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
von: Tang, Shixiang, et al.
Veröffentlicht: (2025)
von: Tang, Shixiang, et al.
Veröffentlicht: (2025)
OmniGAIA: Towards Native Omni-Modal AI Agents
von: Li, Xiaoxi, et al.
Veröffentlicht: (2026)
von: Li, Xiaoxi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
PhysNav-DG: A Novel Adaptive Framework for Robust VLM-Sensor Fusion in Navigation Applications
von: Srinivasan, Trisanth, et al.
Veröffentlicht: (2025) -
RMAdapter: Reconstruction-based Multi-Modal Adapter for Vision-Language Models
von: Lin, Xiang, et al.
Veröffentlicht: (2025) -
Goal-Based Vision-Language Driving
von: Patapati, Santosh, et al.
Veröffentlicht: (2025) -
PlanLLM: Video Procedure Planning with Refinable Large Language Models
von: Yang, Dejie, et al.
Veröffentlicht: (2024) -
Towards Multi-Task Multi-Modal Models: A Video Generative Perspective
von: Yu, Lijun
Veröffentlicht: (2024)