Learning the meanings of function words from grounded language using a visual question answering model
Fuente:
arXiv
Guardado en:
| Autores principales: | Portelance, Eva, Frank, Michael C., Jurafsky, Dan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Reframing linguistic bootstrapping as joint inference using visually-grounded grammar induction models
por: Portelance, Eva, et al.
Publicado: (2024)
por: Portelance, Eva, et al.
Publicado: (2024)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
por: Yang, Shan
Publicado: (2026)
por: Yang, Shan
Publicado: (2026)
Survey Transfer Learning: Recycling Data with Silicon Responses
por: Amini, Ali
Publicado: (2025)
por: Amini, Ali
Publicado: (2025)
TensLoRA: Tensor Alternatives for Low-Rank Adaptation
por: Marmoret, Axel, et al.
Publicado: (2025)
por: Marmoret, Axel, et al.
Publicado: (2025)
QoSGMAA: A Robust Multi-Order Graph Attention and Adversarial Framework for Sparse QoS Prediction
por: Du, Guanchen, et al.
Publicado: (2025)
por: Du, Guanchen, et al.
Publicado: (2025)
SALLIE: Safeguarding Against Latent Language & Image Exploits
por: Azov, Guy, et al.
Publicado: (2026)
por: Azov, Guy, et al.
Publicado: (2026)
Unpacking Failure Modes of Generative Policies: Runtime Monitoring of Consistency and Progress
por: Agia, Christopher, et al.
Publicado: (2024)
por: Agia, Christopher, et al.
Publicado: (2024)
Analyzing Quality, Bias, and Performance in Text-to-Image Generative Models
por: Masrourisaadat, Nila, et al.
Publicado: (2024)
por: Masrourisaadat, Nila, et al.
Publicado: (2024)
On the Compatibility of Generative AI and Generative Linguistics
por: Portelance, Eva, et al.
Publicado: (2024)
por: Portelance, Eva, et al.
Publicado: (2024)
Deployment-Time Reliability of Learned Robot Policies
por: Agia, Christopher
Publicado: (2026)
por: Agia, Christopher
Publicado: (2026)
Aligning by Misaligning: Boundary-aware Curriculum Learning for Multimodal Alignment
por: Ye, Hua, et al.
Publicado: (2025)
por: Ye, Hua, et al.
Publicado: (2025)
WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents
por: Liu, Bingnan, et al.
Publicado: (2026)
por: Liu, Bingnan, et al.
Publicado: (2026)
SpatialMath: Spatial Comprehension-Infused Symbolic Reasoning for Mathematical Problem-Solving
por: Bajpai, Ashutosh, et al.
Publicado: (2026)
por: Bajpai, Ashutosh, et al.
Publicado: (2026)
Waking Up Blind: Cold-Start Optimization of Supervision-Free Agentic Trajectories for Grounded Visual Perception
por: Bajpai, Ashutosh, et al.
Publicado: (2026)
por: Bajpai, Ashutosh, et al.
Publicado: (2026)
A Roadmap for Multilingual, Multimodal Domain Independent Deception Detection
por: Boumber, Dainis, et al.
Publicado: (2024)
por: Boumber, Dainis, et al.
Publicado: (2024)
Text-to-Events: Synthetic Event Camera Streams from Conditional Text Input
por: Ott, Joachim, et al.
Publicado: (2024)
por: Ott, Joachim, et al.
Publicado: (2024)
Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
por: Tu, Songjun, et al.
Publicado: (2025)
por: Tu, Songjun, et al.
Publicado: (2025)
GLoT: A Novel Gated-Logarithmic Transformer for Efficient Sign Language Translation
por: Shahin, Nada, et al.
Publicado: (2025)
por: Shahin, Nada, et al.
Publicado: (2025)
Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight
por: Romero, Angel, et al.
Publicado: (2025)
por: Romero, Angel, et al.
Publicado: (2025)
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
por: Li, Danyang, et al.
Publicado: (2025)
por: Li, Danyang, et al.
Publicado: (2025)
Taking Flight with Dialogue: Enabling Natural Language Control for PX4-based Drone Agent
por: Lim, Shoon Kit, et al.
Publicado: (2025)
por: Lim, Shoon Kit, et al.
Publicado: (2025)
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024)
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024)
SUN Team's Contribution to ABAW 2024 Competition: Audio-visual Valence-Arousal Estimation and Expression Recognition
por: Dresvyanskiy, Denis, et al.
Publicado: (2024)
por: Dresvyanskiy, Denis, et al.
Publicado: (2024)
VLM-VPI: A Vision-Language Reasoning Framework for Improving Automated Vehicle-Pedestrian Interactions
por: Pu, Qingwen, et al.
Publicado: (2026)
por: Pu, Qingwen, et al.
Publicado: (2026)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
por: Dai, Song, et al.
Publicado: (2025)
por: Dai, Song, et al.
Publicado: (2025)
Universal Adversarial Attack on Aligned Multimodal LLMs
por: Rahmatullaev, Temurbek, et al.
Publicado: (2025)
por: Rahmatullaev, Temurbek, et al.
Publicado: (2025)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
por: Tong, Jingqi, et al.
Publicado: (2025)
por: Tong, Jingqi, et al.
Publicado: (2025)
Memory-Efficient Differentially Private Training with Gradient Random Projection
por: Mulrooney, Alex, et al.
Publicado: (2025)
por: Mulrooney, Alex, et al.
Publicado: (2025)
PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model
por: Zhang, Sinin, et al.
Publicado: (2026)
por: Zhang, Sinin, et al.
Publicado: (2026)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
por: Bian, Zhipeng, et al.
Publicado: (2026)
por: Bian, Zhipeng, et al.
Publicado: (2026)
AGOP as Explanation: From Feature Learning to Per-Sample Attribution in Image Classifiers
por: Katakam, Raj Kiran Gupta
Publicado: (2026)
por: Katakam, Raj Kiran Gupta
Publicado: (2026)
Unpacking Hateful Memes: Presupposed Context and False Claims
por: Cai, Weibin, et al.
Publicado: (2025)
por: Cai, Weibin, et al.
Publicado: (2025)
WayFASTER: a Self-Supervised Traversability Prediction for Increased Navigation Awareness
por: Gasparino, Mateus Valverde, et al.
Publicado: (2024)
por: Gasparino, Mateus Valverde, et al.
Publicado: (2024)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
por: Fadli, Samih
Publicado: (2025)
por: Fadli, Samih
Publicado: (2025)
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
por: Durrani, Hamza Ahmed, et al.
Publicado: (2026)
por: Durrani, Hamza Ahmed, et al.
Publicado: (2026)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024)
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024)
EduFlow: Advancing MLLMs' Problem-Solving Proficiency through Multi-Stage, Multi-Perspective Critique
por: Zhu, Chenglin, et al.
Publicado: (2025)
por: Zhu, Chenglin, et al.
Publicado: (2025)
Think, Act, Learn: A Framework for Autonomous Robotic Agents using Closed-Loop Large Language Models
por: Menon, Anjali R., et al.
Publicado: (2025)
por: Menon, Anjali R., et al.
Publicado: (2025)
ADAT: Time-Series-Aware Adaptive Transformer Architecture for Sign Language Translation
por: Shahin, Nada, et al.
Publicado: (2025)
por: Shahin, Nada, et al.
Publicado: (2025)
Can phones, syllables, and words emerge as side-products of cross-situational audiovisual learning? -- A computational investigation
por: Khorrami, Khazar, et al.
Publicado: (2021)
por: Khorrami, Khazar, et al.
Publicado: (2021)
Ejemplares similares
-
Reframing linguistic bootstrapping as joint inference using visually-grounded grammar induction models
por: Portelance, Eva, et al.
Publicado: (2024) -
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
por: Yang, Shan
Publicado: (2026) -
Survey Transfer Learning: Recycling Data with Silicon Responses
por: Amini, Ali
Publicado: (2025) -
TensLoRA: Tensor Alternatives for Low-Rank Adaptation
por: Marmoret, Axel, et al.
Publicado: (2025) -
QoSGMAA: A Robust Multi-Order Graph Attention and Adversarial Framework for Sparse QoS Prediction
por: Du, Guanchen, et al.
Publicado: (2025)