Manimator: Transforming Research Papers into Visual Explanations
Fuente:
arXiv
Guardado en:
| Autores principales: | P, Samarth, Jain, Vyoman, Golugula, Shiva, Sathvik, Motamarri Sai |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
From Query to Explanation: Uni-RAG for Multi-Modal Retrieval-Augmented Learning in STEM
por: Wu, Xinyi, et al.
Publicado: (2025)
por: Wu, Xinyi, et al.
Publicado: (2025)
TCC-Bench: Benchmarking the Traditional Chinese Culture Understanding Capabilities of MLLMs
por: Xu, Pengju, et al.
Publicado: (2025)
por: Xu, Pengju, et al.
Publicado: (2025)
A Shift In Artistic Practices through Artificial Intelligence
por: Tatar, Kıvanç, et al.
Publicado: (2023)
por: Tatar, Kıvanç, et al.
Publicado: (2023)
SimInterview: Transforming Business Education through Large Language Model-Based Simulated Multilingual Interview Training System
por: Nguyen, Truong Thanh Hung, et al.
Publicado: (2025)
por: Nguyen, Truong Thanh Hung, et al.
Publicado: (2025)
AI-based System for Transforming text and sound to Educational Videos
por: ElAlami, M. E., et al.
Publicado: (2026)
por: ElAlami, M. E., et al.
Publicado: (2026)
AI Slop or AI-enhancement? Student perceptions of AI-generated media for an English for Academic Purposes course
por: Woo, David James, et al.
Publicado: (2026)
por: Woo, David James, et al.
Publicado: (2026)
ObjFormer: Learning Land-Cover Changes From Paired OSM Data and Optical High-Resolution Imagery via Object-Guided Transformer
por: Chen, Hongruixuan, et al.
Publicado: (2023)
por: Chen, Hongruixuan, et al.
Publicado: (2023)
Designing Singing Syllabi with Virtual Avatars: AI-Assisted Syllabus Reauthoring
por: Wu, Xinxing
Publicado: (2025)
por: Wu, Xinxing
Publicado: (2025)
PediatricsMQA: a Multi-modal Pediatrics Question Answering Benchmark
por: Bahaj, Adil, et al.
Publicado: (2025)
por: Bahaj, Adil, et al.
Publicado: (2025)
Digital Simulations to Enhance Military Medical Evacuation Decision-Making
por: Fischer, Jeremy, et al.
Publicado: (2025)
por: Fischer, Jeremy, et al.
Publicado: (2025)
FeedQUAC: Quick Unobtrusive AI-Generated Commentary
por: Long, Tao, et al.
Publicado: (2025)
por: Long, Tao, et al.
Publicado: (2025)
MORALISE: A Structured Benchmark for Moral Alignment in Visual Language Models
por: Lin, Xiao, et al.
Publicado: (2025)
por: Lin, Xiao, et al.
Publicado: (2025)
Integration of Policy and Reputation based Trust Mechanisms in e-Commerce Industry
por: Siddiqui, Muhammad Yasir, et al.
Publicado: (2024)
por: Siddiqui, Muhammad Yasir, et al.
Publicado: (2024)
Can LLMs Create Legally Relevant Summaries and Analyses of Videos?
por: Hoeben-Kuil, Lyra, et al.
Publicado: (2025)
por: Hoeben-Kuil, Lyra, et al.
Publicado: (2025)
Bridging the Data Provenance Gap Across Text, Speech and Video
por: Longpre, Shayne, et al.
Publicado: (2024)
por: Longpre, Shayne, et al.
Publicado: (2024)
Harmful YouTube Video Detection: A Taxonomy of Online Harm and MLLMs as Alternative Annotators
por: Jo, Claire Wonjeong, et al.
Publicado: (2024)
por: Jo, Claire Wonjeong, et al.
Publicado: (2024)
KI-Bilder und die Widerständigkeit der Medienkonvergenz: Von primärer zu sekundärer Intermedialität?
por: Wilde, Lukas R. A.
Publicado: (2024)
por: Wilde, Lukas R. A.
Publicado: (2024)
Towards nation-wide analytical healthcare infrastructures: A privacy-preserving augmented knee rehabilitation case study
por: Bačić, Boris, et al.
Publicado: (2024)
por: Bačić, Boris, et al.
Publicado: (2024)
TraceRouter: Robust Safety for Large Foundation Models via Path-Level Intervention
por: Shi, Chuancheng, et al.
Publicado: (2026)
por: Shi, Chuancheng, et al.
Publicado: (2026)
SlideTailor: Personalized Presentation Slide Generation for Scientific Papers
por: Zeng, Wenzheng, et al.
Publicado: (2025)
por: Zeng, Wenzheng, et al.
Publicado: (2025)
Talking Slide Avatars: Open-Source Multimodal Communication Approach for Teaching
por: Wu, Xinxing
Publicado: (2026)
por: Wu, Xinxing
Publicado: (2026)
Bridging the behavior-neural gap: A multimodal AI reveals the brain's geometry of emotion more accurately than human self-reports
por: Du, Changde, et al.
Publicado: (2025)
por: Du, Changde, et al.
Publicado: (2025)
See, Think, Act: Online Shopper Behavior Simulation with VLM Agents
por: Zhang, Yimeng, et al.
Publicado: (2025)
por: Zhang, Yimeng, et al.
Publicado: (2025)
Next-Gen Education: Enhancing AI for Microlearning
por: Saha, Suman, et al.
Publicado: (2025)
por: Saha, Suman, et al.
Publicado: (2025)
Learning to Mask and Permute Visual Tokens for Vision Transformer Pre-Training
por: Baraldi, Lorenzo, et al.
Publicado: (2023)
por: Baraldi, Lorenzo, et al.
Publicado: (2023)
Neuron Abandoning Attention Flow: Visual Explanation of Dynamics inside CNN Models
por: Liao, Yi, et al.
Publicado: (2024)
por: Liao, Yi, et al.
Publicado: (2024)
History-Guided Iterative Visual Reasoning with Self-Correction
por: Yang, Xinglong, et al.
Publicado: (2026)
por: Yang, Xinglong, et al.
Publicado: (2026)
Towards Robust Evaluation of STEM Education: Leveraging MLLMs in Project-Based Learning
por: Wu, Xinyi, et al.
Publicado: (2025)
por: Wu, Xinyi, et al.
Publicado: (2025)
SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection
por: Qi, Peng, et al.
Publicado: (2024)
por: Qi, Peng, et al.
Publicado: (2024)
ANVIL: Analogies and Videos for Lecturers
por: Noviello, Yuri, et al.
Publicado: (2026)
por: Noviello, Yuri, et al.
Publicado: (2026)
GeoGuess: Multimodal Reasoning based on Hierarchy of Visual Information in Street View
por: Cheng, Fenghua, et al.
Publicado: (2025)
por: Cheng, Fenghua, et al.
Publicado: (2025)
Enhancing Student Feedback Using Predictive Models in Visual Literacy Courses
por: Friedman, Alon, et al.
Publicado: (2024)
por: Friedman, Alon, et al.
Publicado: (2024)
Embedding an Ethical Mind: Aligning Text-to-Image Synthesis via Lightweight Value Optimization
por: Wang, Xingqi, et al.
Publicado: (2024)
por: Wang, Xingqi, et al.
Publicado: (2024)
Unmasking Illusions: Understanding Human Perception of Audiovisual Deepfakes
por: Hashmi, Ammarah, et al.
Publicado: (2024)
por: Hashmi, Ammarah, et al.
Publicado: (2024)
M3DR: Towards Universal Multilingual Multimodal Document Retrieval
por: Kolavi, Adithya S, et al.
Publicado: (2025)
por: Kolavi, Adithya S, et al.
Publicado: (2025)
Differential Multimodal Transformers
por: Li, Jerry, et al.
Publicado: (2025)
por: Li, Jerry, et al.
Publicado: (2025)
AUREXA-SE: Audio-Visual Unified Representation Exchange Architecture with Cross-Attention and Squeezeformer for Speech Enhancement
por: Sajid, M., et al.
Publicado: (2025)
por: Sajid, M., et al.
Publicado: (2025)
Open-Vocabulary Audio-Visual Semantic Segmentation
por: Guo, Ruohao, et al.
Publicado: (2024)
por: Guo, Ruohao, et al.
Publicado: (2024)
LLM2Manim: Pedagogy-Aware AI Generation of STEM Animations
por: Joshi, Aastha, et al.
Publicado: (2026)
por: Joshi, Aastha, et al.
Publicado: (2026)
The heteronomy of algorithms: Traditional knowledge and computational knowledge
por: Berry, David M.
Publicado: (2025)
por: Berry, David M.
Publicado: (2025)
Ejemplares similares
-
From Query to Explanation: Uni-RAG for Multi-Modal Retrieval-Augmented Learning in STEM
por: Wu, Xinyi, et al.
Publicado: (2025) -
TCC-Bench: Benchmarking the Traditional Chinese Culture Understanding Capabilities of MLLMs
por: Xu, Pengju, et al.
Publicado: (2025) -
A Shift In Artistic Practices through Artificial Intelligence
por: Tatar, Kıvanç, et al.
Publicado: (2023) -
SimInterview: Transforming Business Education through Large Language Model-Based Simulated Multilingual Interview Training System
por: Nguyen, Truong Thanh Hung, et al.
Publicado: (2025) -
AI-based System for Transforming text and sound to Educational Videos
por: ElAlami, M. E., et al.
Publicado: (2026)