Collaborative AI Enhances Image Understanding in Materials Science
Fuente:
arXiv
Saved in:
| Main Authors: | Yin, Ruoyan Avery, Ren, Zhichu, Yin, Zongyou, Zhang, Zhen, Kim, So Yeon, Hsu, Chia-Wei, Li, Ju |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OmniAcc: Personalized Accessibility Assistant Using Generative AI
by: Karki, Siddhant, et al.
Published: (2025)
by: Karki, Siddhant, et al.
Published: (2025)
SUN Team's Contribution to ABAW 2024 Competition: Audio-visual Valence-Arousal Estimation and Expression Recognition
by: Dresvyanskiy, Denis, et al.
Published: (2024)
by: Dresvyanskiy, Denis, et al.
Published: (2024)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
by: Raoufi, Behnam, et al.
Published: (2025)
by: Raoufi, Behnam, et al.
Published: (2025)
Lightweight Low-SNR-Robust Semantic Communication System for Autonomous Driving
by: Ren, Ruixing, et al.
Published: (2026)
by: Ren, Ruixing, et al.
Published: (2026)
AI-Dentify: Deep learning for proximal caries detection on bitewing x-ray -- HUNT4 Oral Health Study
by: de Frutos, Javier Pérez, et al.
Published: (2023)
by: de Frutos, Javier Pérez, et al.
Published: (2023)
Vectra: A New Metric, Dataset, and Model for Visual Quality Assessment in E-Commerce In-Image Machine Translation
by: Wu, Qingyu, et al.
Published: (2026)
by: Wu, Qingyu, et al.
Published: (2026)
U-Net-Like Spiking Neural Networks for Single Image Dehazing
by: Li, Huibin, et al.
Published: (2025)
by: Li, Huibin, et al.
Published: (2025)
Ultra-Reduced-Impact-Encased-Logging (URIEL): propose a new method for selective sustainable logging and post-harvest silvicultural treatment in tropical forest using airborne robotics systems
by: Albiero, Daniel, et al.
Published: (2026)
by: Albiero, Daniel, et al.
Published: (2026)
FlightScope: An Experimental Comparative Review of Aircraft Detection Algorithms in Satellite Imagery
by: Ghazouali, Safouane El, et al.
Published: (2024)
by: Ghazouali, Safouane El, et al.
Published: (2024)
Deep Learning methodology for the identification of wood species using high-resolution macroscopic images
by: Herrera-Poyatos, David, et al.
Published: (2024)
by: Herrera-Poyatos, David, et al.
Published: (2024)
Habitat Classification from Ground-Level Imagery Using Deep Neural Networks
by: Shi, Hongrui, et al.
Published: (2025)
by: Shi, Hongrui, et al.
Published: (2025)
Precision at Scale: Domain-Specific Datasets On-Demand
by: Rodríguez-de-Vera, Jesús M, et al.
Published: (2024)
by: Rodríguez-de-Vera, Jesús M, et al.
Published: (2024)
A Segmented Robot Grasping Perception Neural Network for Edge AI
by: Bröcheler, Casper, et al.
Published: (2025)
by: Bröcheler, Casper, et al.
Published: (2025)
Open Gaze: Open Source eye tracker for smartphone devices using Deep Learning
by: reddy, Sushmanth, et al.
Published: (2023)
by: reddy, Sushmanth, et al.
Published: (2023)
AGOP as Explanation: From Feature Learning to Per-Sample Attribution in Image Classifiers
by: Katakam, Raj Kiran Gupta
Published: (2026)
by: Katakam, Raj Kiran Gupta
Published: (2026)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
by: Bian, Zhipeng, et al.
Published: (2026)
by: Bian, Zhipeng, et al.
Published: (2026)
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
by: Durrani, Hamza Ahmed, et al.
Published: (2026)
by: Durrani, Hamza Ahmed, et al.
Published: (2026)
Motion Perceiver: Real-Time Occupancy Forecasting for Embedded Systems
by: Ferenczi, Bryce, et al.
Published: (2023)
by: Ferenczi, Bryce, et al.
Published: (2023)
Towards Cognitive Collaborative Robots: Semantic-Level Integration and Explainable Control for Human-Centric Cooperation
by: Oh, Jaehong
Published: (2025)
by: Oh, Jaehong
Published: (2025)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
by: Tong, Jingqi, et al.
Published: (2025)
by: Tong, Jingqi, et al.
Published: (2025)
CCVA-FL: Cross-Client Variations Adaptive Federated Learning for Medical Imaging
by: Gupta, Sunny, et al.
Published: (2024)
by: Gupta, Sunny, et al.
Published: (2024)
PyCAT4: A Hierarchical Vision Transformer-based Framework for 3D Human Pose Estimation
by: Yang, Zongyou, et al.
Published: (2025)
by: Yang, Zongyou, et al.
Published: (2025)
CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment
by: Kang, Li, et al.
Published: (2026)
by: Kang, Li, et al.
Published: (2026)
H-FCBFormer Hierarchical Fully Convolutional Branch Transformer for Occlusal Contact Segmentation with Articulating Paper
by: Banks, Ryan, et al.
Published: (2024)
by: Banks, Ryan, et al.
Published: (2024)
Taming the Tail: Leveraging Asymmetric Loss and Pade Approximation to Overcome Medical Image Long-Tailed Class Imbalance
by: Kashyap, Pankhi, et al.
Published: (2024)
by: Kashyap, Pankhi, et al.
Published: (2024)
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
by: Li, Danyang, et al.
Published: (2025)
by: Li, Danyang, et al.
Published: (2025)
Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight
by: Romero, Angel, et al.
Published: (2025)
by: Romero, Angel, et al.
Published: (2025)
WayFASTER: a Self-Supervised Traversability Prediction for Increased Navigation Awareness
by: Gasparino, Mateus Valverde, et al.
Published: (2024)
by: Gasparino, Mateus Valverde, et al.
Published: (2024)
EmbodiedLGR: Integrating Lightweight Graph Representation and Retrieval for Semantic-Spatial Memory in Robotic Agents
by: Riva, Paolo, et al.
Published: (2026)
by: Riva, Paolo, et al.
Published: (2026)
Taking Flight with Dialogue: Enabling Natural Language Control for PX4-based Drone Agent
by: Lim, Shoon Kit, et al.
Published: (2025)
by: Lim, Shoon Kit, et al.
Published: (2025)
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
by: Gopinathan, Muraleekrishna, et al.
Published: (2024)
by: Gopinathan, Muraleekrishna, et al.
Published: (2024)
Deep Probabilistic Traversability with Test-time Adaptation for Uncertainty-aware Planetary Rover Navigation
by: Endo, Masafumi, et al.
Published: (2024)
by: Endo, Masafumi, et al.
Published: (2024)
CoMoCAVs: Cohesive Decision-Guided Motion Planning for Connected and Autonomous Vehicles with Multi-Policy Reinforcement Learning
by: Hu, Pan
Published: (2025)
by: Hu, Pan
Published: (2025)
NeuGrasp: Generalizable Neural Surface Reconstruction with Background Priors for Material-Agnostic Object Grasp Detection
by: Fan, Qingyu, et al.
Published: (2025)
by: Fan, Qingyu, et al.
Published: (2025)
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
by: Anh, Duy Le Dinh, et al.
Published: (2024)
by: Anh, Duy Le Dinh, et al.
Published: (2024)
Learning Association via Track-Detection Matching for Multi-Object Tracking
by: Adžemović, Momir
Published: (2025)
by: Adžemović, Momir
Published: (2025)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
by: Dai, Song, et al.
Published: (2025)
by: Dai, Song, et al.
Published: (2025)
Universal Adversarial Attack on Aligned Multimodal LLMs
by: Rahmatullaev, Temurbek, et al.
Published: (2025)
by: Rahmatullaev, Temurbek, et al.
Published: (2025)
SoccerRef-Agents: Multi-Agent System for Automated Soccer Refereeing
by: Meng, Zi, et al.
Published: (2026)
by: Meng, Zi, et al.
Published: (2026)
Memory-Efficient Differentially Private Training with Gradient Random Projection
by: Mulrooney, Alex, et al.
Published: (2025)
by: Mulrooney, Alex, et al.
Published: (2025)
Similar Items
-
OmniAcc: Personalized Accessibility Assistant Using Generative AI
by: Karki, Siddhant, et al.
Published: (2025) -
SUN Team's Contribution to ABAW 2024 Competition: Audio-visual Valence-Arousal Estimation and Expression Recognition
by: Dresvyanskiy, Denis, et al.
Published: (2024) -
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
by: Raoufi, Behnam, et al.
Published: (2025) -
Lightweight Low-SNR-Robust Semantic Communication System for Autonomous Driving
by: Ren, Ruixing, et al.
Published: (2026) -
AI-Dentify: Deep learning for proximal caries detection on bitewing x-ray -- HUNT4 Oral Health Study
by: de Frutos, Javier Pérez, et al.
Published: (2023)