ViSTa Dataset: Do vision-language models understand sequential tasks?
Fuente:
arXiv
Guardado en:
| Autores principales: | Wybitul, Evžen, Gunter, Evan Ryan, Seleznyov, Mikhail, Lindner, David |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Representations of Text and Images Align From Layer One
por: Wybitul, Evžen, et al.
Publicado: (2026)
por: Wybitul, Evžen, et al.
Publicado: (2026)
Amortizing intractable inference in diffusion models for vision, language, and control
por: Venkatraman, Siddarth, et al.
Publicado: (2024)
por: Venkatraman, Siddarth, et al.
Publicado: (2024)
Do large language vision models understand 3D shapes?
por: Eppel, Sagi
Publicado: (2024)
por: Eppel, Sagi
Publicado: (2024)
The Solution for the sequential task continual learning track of the 2nd Greater Bay Area International Algorithm Competition
por: Pan, Sishun, et al.
Publicado: (2024)
por: Pan, Sishun, et al.
Publicado: (2024)
STaTS: Structure-Aware Temporal Sequence Summarization via Statistical Window Merging
por: Bhowmick, Disharee, et al.
Publicado: (2025)
por: Bhowmick, Disharee, et al.
Publicado: (2025)
Visual hallucination detection in large vision-language models via evidential conflict
por: Huang, Tao, et al.
Publicado: (2025)
por: Huang, Tao, et al.
Publicado: (2025)
FlySearch: Exploring how vision-language models explore
por: Pardyl, Adam, et al.
Publicado: (2025)
por: Pardyl, Adam, et al.
Publicado: (2025)
What do vision-language models see in the context? Investigating multimodal in-context learning
por: Santos, Gabriel O. dos, et al.
Publicado: (2025)
por: Santos, Gabriel O. dos, et al.
Publicado: (2025)
Contrastive vision-language learning with paraphrasing and negation
por: Ngan, Kwun Ho, et al.
Publicado: (2025)
por: Ngan, Kwun Ho, et al.
Publicado: (2025)
Comprehensive language-image pre-training for 3D medical image understanding
por: Wald, Tassilo, et al.
Publicado: (2025)
por: Wald, Tassilo, et al.
Publicado: (2025)
BRAVE: Broadening the visual encoding of vision-language models
por: Kar, Oğuzhan Fatih, et al.
Publicado: (2024)
por: Kar, Oğuzhan Fatih, et al.
Publicado: (2024)
Do generative video models understand physical principles?
por: Motamed, Saman, et al.
Publicado: (2025)
por: Motamed, Saman, et al.
Publicado: (2025)
Efficient training for compact compression models via sequential distillation
por: Rodrigues, Caroline Mazini, et al.
Publicado: (2026)
por: Rodrigues, Caroline Mazini, et al.
Publicado: (2026)
Video Annotator: A framework for efficiently building video classifiers using vision-language models and active learning
por: Ziai, Amir, et al.
Publicado: (2024)
por: Ziai, Amir, et al.
Publicado: (2024)
ViTCAE: ViT-based Class-conditioned Autoencoder
por: Jebraeeli, Vahid, et al.
Publicado: (2025)
por: Jebraeeli, Vahid, et al.
Publicado: (2025)
The in-context inductive biases of vision-language models differ across modalities
por: Allen, Kelsey, et al.
Publicado: (2025)
por: Allen, Kelsey, et al.
Publicado: (2025)
GP-VLS: A general-purpose vision language model for surgery
por: Schmidgall, Samuel, et al.
Publicado: (2024)
por: Schmidgall, Samuel, et al.
Publicado: (2024)
ProtoS-ViT: Visual foundation models for sparse self-explainable classifications
por: Turbé, Hugues, et al.
Publicado: (2024)
por: Turbé, Hugues, et al.
Publicado: (2024)
STRAP-ViT: Segregated Tokens with Randomized -- Transformations for Defense against Adversarial Patches in ViTs
por: Chattopadhyay, Nandish, et al.
Publicado: (2026)
por: Chattopadhyay, Nandish, et al.
Publicado: (2026)
Representation geometry shapes task performance in vision-language modeling for CT enterography
por: Minoccheri, Cristian, et al.
Publicado: (2026)
por: Minoccheri, Cristian, et al.
Publicado: (2026)
Cross-modal linkage risk in clinical vision-language models
por: Arasteh, Soroosh Tayebi, et al.
Publicado: (2026)
por: Arasteh, Soroosh Tayebi, et al.
Publicado: (2026)
PathAlign: A vision-language model for whole slide images in histopathology
por: Ahmed, Faruk, et al.
Publicado: (2024)
por: Ahmed, Faruk, et al.
Publicado: (2024)
TRecViT: A Recurrent Video Transformer
por: Pătrăucean, Viorica, et al.
Publicado: (2024)
por: Pătrăucean, Viorica, et al.
Publicado: (2024)
How to train your ViT for OOD Detection
por: Mueller, Maximilian, et al.
Publicado: (2024)
por: Mueller, Maximilian, et al.
Publicado: (2024)
Linking heterogeneous microstructure informatics with expert characterization knowledge through customized and hybrid vision-language representations for industrial qualification
por: Safdar, Mutahar, et al.
Publicado: (2025)
por: Safdar, Mutahar, et al.
Publicado: (2025)
MC-ViViT: Multi-branch Classifier-ViViT to detect Mild Cognitive Impairment in older adults using facial videos
por: Sun, Jian, et al.
Publicado: (2023)
por: Sun, Jian, et al.
Publicado: (2023)
ViViDex: Learning Vision-based Dexterous Manipulation from Human Videos
por: Chen, Zerui, et al.
Publicado: (2024)
por: Chen, Zerui, et al.
Publicado: (2024)
HydraViT: Stacking Heads for a Scalable ViT
por: Haberer, Janek, et al.
Publicado: (2024)
por: Haberer, Janek, et al.
Publicado: (2024)
WriteViT: Handwritten Text Generation with Vision Transformer
por: Nam, Dang Hoai, et al.
Publicado: (2025)
por: Nam, Dang Hoai, et al.
Publicado: (2025)
Look, Remember and Reason: Grounded reasoning in videos with language models
por: Bhattacharyya, Apratim, et al.
Publicado: (2023)
por: Bhattacharyya, Apratim, et al.
Publicado: (2023)
Advancing vision-language models in front-end development via data synthesis
por: Ge, Tong, et al.
Publicado: (2025)
por: Ge, Tong, et al.
Publicado: (2025)
Intriguing Frequency Interpretation of Adversarial Robustness for CNNs and ViTs
por: Chen, Lu, et al.
Publicado: (2025)
por: Chen, Lu, et al.
Publicado: (2025)
Exploiting Lightweight Hierarchical ViT and Dynamic Framework for Efficient Visual Tracking
por: Kang, Ben, et al.
Publicado: (2025)
por: Kang, Ben, et al.
Publicado: (2025)
Token Cropr: Faster ViTs for Quite a Few Tasks
por: Bergner, Benjamin, et al.
Publicado: (2024)
por: Bergner, Benjamin, et al.
Publicado: (2024)
LD-ViCE: Latent Diffusion Model for Video Counterfactual Explanations
por: Varshney, Payal, et al.
Publicado: (2025)
por: Varshney, Payal, et al.
Publicado: (2025)
CubistMerge: Spatial-Preserving Token Merging For Diverse ViT Backbones
por: Gong, Wenyi, et al.
Publicado: (2025)
por: Gong, Wenyi, et al.
Publicado: (2025)
LongViTU: Instruction Tuning for Long-Form Video Understanding
por: Wu, Rujie, et al.
Publicado: (2025)
por: Wu, Rujie, et al.
Publicado: (2025)
Medical Semantic Segmentation with Diffusion Pretrain
por: Li, David, et al.
Publicado: (2025)
por: Li, David, et al.
Publicado: (2025)
Decoupling the components of geometric understanding in Vision Language Models
por: Kosoy, Eliza, et al.
Publicado: (2025)
por: Kosoy, Eliza, et al.
Publicado: (2025)
VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions
por: Bulat, Adrian, et al.
Publicado: (2026)
por: Bulat, Adrian, et al.
Publicado: (2026)
Ejemplares similares
-
Representations of Text and Images Align From Layer One
por: Wybitul, Evžen, et al.
Publicado: (2026) -
Amortizing intractable inference in diffusion models for vision, language, and control
por: Venkatraman, Siddarth, et al.
Publicado: (2024) -
Do large language vision models understand 3D shapes?
por: Eppel, Sagi
Publicado: (2024) -
The Solution for the sequential task continual learning track of the 2nd Greater Bay Area International Algorithm Competition
por: Pan, Sishun, et al.
Publicado: (2024) -
STaTS: Structure-Aware Temporal Sequence Summarization via Statistical Window Merging
por: Bhowmick, Disharee, et al.
Publicado: (2025)