A multi-purpose automatic editing system based on lecture semantics for remote education
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Panwen, Huang, Rui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
Cross Modification Attention Based Deliberation Model for Image Captioning
von: Lian, Zheng, et al.
Veröffentlicht: (2021)
von: Lian, Zheng, et al.
Veröffentlicht: (2021)
Case-based reasoning approach for diagnostic screening of children with developmental delays
von: Song, Zichen, et al.
Veröffentlicht: (2024)
von: Song, Zichen, et al.
Veröffentlicht: (2024)
Terrain Diffusion Network: Climatic-Aware Terrain Generation with Geological Sketch Guidance
von: Hu, Zexin, et al.
Veröffentlicht: (2023)
von: Hu, Zexin, et al.
Veröffentlicht: (2023)
Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering
von: Zhang, Yanjie, et al.
Veröffentlicht: (2026)
von: Zhang, Yanjie, et al.
Veröffentlicht: (2026)
Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment
von: Cai, Zhuoxuan, et al.
Veröffentlicht: (2025)
von: Cai, Zhuoxuan, et al.
Veröffentlicht: (2025)
Boosting Audio Visual Question Answering via Key Semantic-Aware Cues
von: Li, Guangyao, et al.
Veröffentlicht: (2024)
von: Li, Guangyao, et al.
Veröffentlicht: (2024)
Anti-Inpainting: A Proactive Defense Approach against Malicious Diffusion-based Inpainters under Unknown Conditions
von: Guo, Yimao, et al.
Veröffentlicht: (2025)
von: Guo, Yimao, et al.
Veröffentlicht: (2025)
Exploiting LMM-based knowledge for image classification tasks
von: Tzelepi, Maria, et al.
Veröffentlicht: (2024)
von: Tzelepi, Maria, et al.
Veröffentlicht: (2024)
Diagnosing and Re-learning for Balanced Multimodal Learning
von: Wei, Yake, et al.
Veröffentlicht: (2024)
von: Wei, Yake, et al.
Veröffentlicht: (2024)
Webcam-based Pupil Diameter Prediction Benefits from Upscaling
von: Shah, Vijul, et al.
Veröffentlicht: (2024)
von: Shah, Vijul, et al.
Veröffentlicht: (2024)
KAN Text to Vision? The Exploration of Kolmogorov-Arnold Networks for Multi-Scale Sequence-Based Pose Animation from Sign Language Notation
von: Du, Guanyi, et al.
Veröffentlicht: (2026)
von: Du, Guanyi, et al.
Veröffentlicht: (2026)
LPM 1.0: Video-based Character Performance Model
von: Zeng, Ailing, et al.
Veröffentlicht: (2026)
von: Zeng, Ailing, et al.
Veröffentlicht: (2026)
Crab$^{+}$: A Scalable and Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
von: Cai, Dongnuan, et al.
Veröffentlicht: (2026)
von: Cai, Dongnuan, et al.
Veröffentlicht: (2026)
PROVE: A Perceptual RemOVal cohErence Benchmark for Visual Media
von: Li, Fuhao, et al.
Veröffentlicht: (2026)
von: Li, Fuhao, et al.
Veröffentlicht: (2026)
LLM-based Fusion of Multi-modal Features for Commercial Memorability Prediction
von: Pramov, Aleksandar
Veröffentlicht: (2025)
von: Pramov, Aleksandar
Veröffentlicht: (2025)
Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
von: Liu, Jiajun, et al.
Veröffentlicht: (2024)
von: Liu, Jiajun, et al.
Veröffentlicht: (2024)
ReCorD: Reasoning and Correcting Diffusion for HOI Generation
von: Jiang-Lin, Jian-Yu, et al.
Veröffentlicht: (2024)
von: Jiang-Lin, Jian-Yu, et al.
Veröffentlicht: (2024)
PFB-Diff: Progressive Feature Blending Diffusion for Text-driven Image Editing
von: Huang, Wenjing, et al.
Veröffentlicht: (2023)
von: Huang, Wenjing, et al.
Veröffentlicht: (2023)
InstructFLIP: Exploring Unified Vision-Language Model for Face Anti-spoofing
von: Lin, Kun-Hsiang, et al.
Veröffentlicht: (2025)
von: Lin, Kun-Hsiang, et al.
Veröffentlicht: (2025)
Omni-Dish: Photorealistic and Faithful Image Generation and Editing for Arbitrary Chinese Dishes
von: Liu, Huijie, et al.
Veröffentlicht: (2025)
von: Liu, Huijie, et al.
Veröffentlicht: (2025)
Detecting Multimodal Situations with Insufficient Context and Abstaining from Baseless Predictions
von: Liu, Junzhang, et al.
Veröffentlicht: (2024)
von: Liu, Junzhang, et al.
Veröffentlicht: (2024)
X-Adapter: Adding Universal Compatibility of Plugins for Upgraded Diffusion Model
von: Ran, Lingmin, et al.
Veröffentlicht: (2023)
von: Ran, Lingmin, et al.
Veröffentlicht: (2023)
High-fidelity and Lip-synced Talking Face Synthesis via Landmark-based Diffusion Model
von: Zhong, Weizhi, et al.
Veröffentlicht: (2024)
von: Zhong, Weizhi, et al.
Veröffentlicht: (2024)
Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation
von: Huang, Feizhen, et al.
Veröffentlicht: (2025)
von: Huang, Feizhen, et al.
Veröffentlicht: (2025)
Robust Fuzzy Multi-view Learning under View Conflict
von: Duan, Siyuan, et al.
Veröffentlicht: (2026)
von: Duan, Siyuan, et al.
Veröffentlicht: (2026)
Can Large Language Models Grasp Event Signals? Exploring Pure Zero-Shot Event-based Recognition
von: Yu, Zongyou, et al.
Veröffentlicht: (2024)
von: Yu, Zongyou, et al.
Veröffentlicht: (2024)
PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation
von: Ling, Jun, et al.
Veröffentlicht: (2024)
von: Ling, Jun, et al.
Veröffentlicht: (2024)
CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization
von: Zhou, Jinxing, et al.
Veröffentlicht: (2025)
von: Zhou, Jinxing, et al.
Veröffentlicht: (2025)
Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models
von: Yu, Xiaomin, et al.
Veröffentlicht: (2026)
von: Yu, Xiaomin, et al.
Veröffentlicht: (2026)
Image is All You Need to Empower Large-scale Diffusion Models for In-Domain Generation
von: Cao, Pu, et al.
Veröffentlicht: (2023)
von: Cao, Pu, et al.
Veröffentlicht: (2023)
Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation
von: Yan, Xin, et al.
Veröffentlicht: (2024)
von: Yan, Xin, et al.
Veröffentlicht: (2024)
CLCR: Cross-Level Semantic Collaborative Representation for Multimodal Learning
von: Meng, Chunlei, et al.
Veröffentlicht: (2026)
von: Meng, Chunlei, et al.
Veröffentlicht: (2026)
Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminability
von: Li, Chengzhi, et al.
Veröffentlicht: (2025)
von: Li, Chengzhi, et al.
Veröffentlicht: (2025)
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
von: Yu, Jiashuo, et al.
Veröffentlicht: (2025)
von: Yu, Jiashuo, et al.
Veröffentlicht: (2025)
AutoAWG: Adverse Weather Generation with Adaptive Multi-Controls for Automotive Videos
von: Hu, Jiagao, et al.
Veröffentlicht: (2026)
von: Hu, Jiagao, et al.
Veröffentlicht: (2026)
DyRoNet: Dynamic Routing and Low-Rank Adapters for Autonomous Driving Streaming Perception
von: Huang, Xiang, et al.
Veröffentlicht: (2024)
von: Huang, Xiang, et al.
Veröffentlicht: (2024)
Pistachio: Towards Synthetic, Balanced, and Long-Form Video Anomaly Benchmarks
von: Li, Jie, et al.
Veröffentlicht: (2025)
von: Li, Jie, et al.
Veröffentlicht: (2025)
MIRROR: Multi-Modal Pathological Self-Supervised Representation Learning via Modality Alignment and Retention
von: Wang, Tianyi, et al.
Veröffentlicht: (2025)
von: Wang, Tianyi, et al.
Veröffentlicht: (2025)
Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR
von: Zhang, Yulong, et al.
Veröffentlicht: (2025)
von: Zhang, Yulong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech
von: Liu, Rui, et al.
Veröffentlicht: (2024) -
Cross Modification Attention Based Deliberation Model for Image Captioning
von: Lian, Zheng, et al.
Veröffentlicht: (2021) -
Case-based reasoning approach for diagnostic screening of children with developmental delays
von: Song, Zichen, et al.
Veröffentlicht: (2024) -
Terrain Diffusion Network: Climatic-Aware Terrain Generation with Geological Sketch Guidance
von: Hu, Zexin, et al.
Veröffentlicht: (2023) -
Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering
von: Zhang, Yanjie, et al.
Veröffentlicht: (2026)