Fine-Tuning Vision-Language Models for Understanding Current Damage and Scoring Priority with Quality Guard Agent
Fuente:
arXiv
Saved in:
| Main Author: | Yasuno, Takato |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Quantized Vision-Language Models for Damage Assessment: A Comparative Study of LLaVA-1.5-7B Quantization Levels
by: Yasuno, Takato
Published: (2026)
by: Yasuno, Takato
Published: (2026)
Proximity QA: Unleashing the Power of Multi-Modal Large Language Models for Spatial Proximity Analysis
by: Li, Jianing, et al.
Published: (2024)
by: Li, Jianing, et al.
Published: (2024)
Cytoarchitecture in Words: Weakly Supervised Vision-Language Modeling for Human Brain Microscopy
by: Sutton, Matthew, et al.
Published: (2026)
by: Sutton, Matthew, et al.
Published: (2026)
Language Predicts Identity Fusion Across Cultures and Reveals Divergent Pathways to Violence
by: Wright, Devin R., et al.
Published: (2026)
by: Wright, Devin R., et al.
Published: (2026)
Heterogeneous Variational Inference for Markov Degradation Hazard Models: Discretized Mixture with Interpretable Clusters
by: Yasuno, Takato
Published: (2026)
by: Yasuno, Takato
Published: (2026)
Understanding Deterioration Random Effects for Causal Discovery in Infrastructure Management
by: Yasuno, Takato
Published: (2026)
by: Yasuno, Takato
Published: (2026)
Chronic pain patient narratives allow for the estimation of current pain intensity
by: Nunes, Diogo A. P., et al.
Published: (2022)
by: Nunes, Diogo A. P., et al.
Published: (2022)
Automating Clinical Information Retrieval from Finnish Electronic Health Records Using Large Language Models
by: Saukkoriipi, Mikko, et al.
Published: (2026)
by: Saukkoriipi, Mikko, et al.
Published: (2026)
Measuring Robustness of Speech Recognition from MEG Signals Under Distribution Shift
by: Chien, Sheng-You, et al.
Published: (2026)
by: Chien, Sheng-You, et al.
Published: (2026)
RAPTOR-AI for Disaster OODA Loop: Hierarchical Multimodal RAG with Experience-Driven Agentic Decision-Making
by: Yasuno, Takato
Published: (2026)
by: Yasuno, Takato
Published: (2026)
Cognitive Linguistic Identity Fusion Score (CLIFS): A Scalable Cognition-Informed Approach to Quantifying Identity Fusion from Text
by: Wright, Devin R., et al.
Published: (2025)
by: Wright, Devin R., et al.
Published: (2025)
Pipeline and Dataset Generation for Automated Fact-checking in Almost Any Language
by: Drchal, Jan, et al.
Published: (2023)
by: Drchal, Jan, et al.
Published: (2023)
Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning
by: Ji, Binbin, et al.
Published: (2025)
by: Ji, Binbin, et al.
Published: (2025)
Scaling Laws for State Dynamics in Large Language Models
by: Li, Jacob X, et al.
Published: (2025)
by: Li, Jacob X, et al.
Published: (2025)
PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation Dialogues
by: Kajare, Prajwal Vijay, et al.
Published: (2026)
by: Kajare, Prajwal Vijay, et al.
Published: (2026)
propella-1: Multi-Property Document Annotation for LLM Data Curation at Scale
by: Idahl, Maximilian, et al.
Published: (2026)
by: Idahl, Maximilian, et al.
Published: (2026)
LongSumEval: Question-Answering Based Evaluation and Feedback-Driven Refinement for Long Document Summarization
by: Nguyen, Huyen, et al.
Published: (2026)
by: Nguyen, Huyen, et al.
Published: (2026)
Suppressing Domain-Specific Hallucination in Construction LLMs: A Knowledge Graph Foundation for GraphRAG and QLoRA on River and Sediment Control Technical Standards
by: Yasuno, Takato
Published: (2026)
by: Yasuno, Takato
Published: (2026)
CG-TTRL: Context-Guided Test-Time Reinforcement Learning for On-Device Large Language Models
by: Hosseini, Peyman, et al.
Published: (2025)
by: Hosseini, Peyman, et al.
Published: (2025)
Thaka at KSAA-2026 Task 2: Regularized Fine-Tuning for Arabic Speech Diacritization
by: Alamr, Meshal, et al.
Published: (2026)
by: Alamr, Meshal, et al.
Published: (2026)
Talking Tennis: Language Feedback from 3D Biomechanical Action Recognition
by: Dashore, Arushi, et al.
Published: (2025)
by: Dashore, Arushi, et al.
Published: (2025)
Adapting Methods for Domain-Specific Japanese Small LMs: Scale, Architecture, and Quantization
by: Yasuno, Takato
Published: (2026)
by: Yasuno, Takato
Published: (2026)
Modular Deep Learning Framework for Assistive Perception: Gaze, Affect, and Speaker Identification
by: Anchan, Akshit Pramod, et al.
Published: (2025)
by: Anchan, Akshit Pramod, et al.
Published: (2025)
Heterogeneous Graph Importance Scoring and Clustering with Automated LLM-based Interpretation
by: Yasuno, Takato
Published: (2026)
by: Yasuno, Takato
Published: (2026)
Transfer-learning for video classification: Video Swin Transformer on multiple domains
by: Oliveira, Daniel A. P., et al.
Published: (2022)
by: Oliveira, Daniel A. P., et al.
Published: (2022)
Multi-stage Bridge Inspection System: Integrating Foundation Models with Location Anonymization
by: Yasuno, Takato
Published: (2026)
by: Yasuno, Takato
Published: (2026)
Distilled HuBERT for Mobile Speech Emotion Recognition: A Cross-Corpus Validation Study
by: Ismail, Saifelden M.
Published: (2025)
by: Ismail, Saifelden M.
Published: (2025)
Detection of Personal Data in Structured Datasets Using a Large Language Model
by: Ntwali, Albert Agisha, et al.
Published: (2025)
by: Ntwali, Albert Agisha, et al.
Published: (2025)
Improving Cross-Lingual Phonetic Representation of Low-Resource Languages Through Language Similarity Analysis
by: Kim, Minu, et al.
Published: (2025)
by: Kim, Minu, et al.
Published: (2025)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
by: Viveiros, André G., et al.
Published: (2025)
by: Viveiros, André G., et al.
Published: (2025)
Multilingual Multi-Label Emotion Classification at Scale with Synthetic Data
by: Borisov, Vadim
Published: (2026)
by: Borisov, Vadim
Published: (2026)
The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering
by: Pandey, Anupam, et al.
Published: (2025)
by: Pandey, Anupam, et al.
Published: (2025)
FS-DAG: Few Shot Domain Adapting Graph Networks for Visually Rich Document Understanding
by: Agarwal, Amit, et al.
Published: (2025)
by: Agarwal, Amit, et al.
Published: (2025)
Logits-Constrained Framework with RoBERTa for Ancient Chinese NER
by: Hua, Wenjie, et al.
Published: (2025)
by: Hua, Wenjie, et al.
Published: (2025)
INESC-ID @ eRisk 2025: Exploring Fine-Tuned, Similarity-Based, and Prompt-Based Approaches to Depression Symptom Identification
by: Nunes, Diogo A. P., et al.
Published: (2025)
by: Nunes, Diogo A. P., et al.
Published: (2025)
On the development of an AI performance and behavioural measures for teaching and classroom management
by: Niculescu, Andreea I., et al.
Published: (2025)
by: Niculescu, Andreea I., et al.
Published: (2025)
Motion-Based Sign Language Video Summarization using Curvature and Torsion
by: Sartinas, Evangelos G., et al.
Published: (2023)
by: Sartinas, Evangelos G., et al.
Published: (2023)
Real-World En Call Center Transcripts Dataset with PII Redaction
by: Dao, Ha, et al.
Published: (2025)
by: Dao, Ha, et al.
Published: (2025)
Structural Stress and Learned Helplessness in Afghanistan: A Multi-Layer Analysis of the AFSTRESS Dari Corpus
by: Baktash, Jawid Ahmad, et al.
Published: (2026)
by: Baktash, Jawid Ahmad, et al.
Published: (2026)
ParaNoise-SV: Integrated Approach for Noise-Robust Speaker Verification with Parallel Joint Learning of Speech Enhancement and Noise Extraction
by: Kim, Minu, et al.
Published: (2025)
by: Kim, Minu, et al.
Published: (2025)
Similar Items
-
Quantized Vision-Language Models for Damage Assessment: A Comparative Study of LLaVA-1.5-7B Quantization Levels
by: Yasuno, Takato
Published: (2026) -
Proximity QA: Unleashing the Power of Multi-Modal Large Language Models for Spatial Proximity Analysis
by: Li, Jianing, et al.
Published: (2024) -
Cytoarchitecture in Words: Weakly Supervised Vision-Language Modeling for Human Brain Microscopy
by: Sutton, Matthew, et al.
Published: (2026) -
Language Predicts Identity Fusion Across Cultures and Reveals Divergent Pathways to Violence
by: Wright, Devin R., et al.
Published: (2026) -
Heterogeneous Variational Inference for Markov Degradation Hazard Models: Discretized Mixture with Interpretable Clusters
by: Yasuno, Takato
Published: (2026)