Dataset Growth
Fuente:
arXiv
Saved in:
| Main Authors: | Qin, Ziheng, Xu, Zhaopan, Zhou, Yukun, Zheng, Zangwei, Cheng, Zebang, Tang, Hao, Shang, Lei, Sun, Baigui, Peng, Xiaojiang, Timofte, Radu, Yao, Hongxun, Wang, Kai, You, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PEBench: A Fictitious Dataset to Benchmark Machine Unlearning for Multimodal Large Language Models
by: Xu, Zhaopan, et al.
Published: (2025)
by: Xu, Zhaopan, et al.
Published: (2025)
MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification
by: Xu, Zhaopan, et al.
Published: (2025)
by: Xu, Zhaopan, et al.
Published: (2025)
DAM: Dual Active Learning with Multimodal Foundation Model for Source-Free Domain Adaptation
by: Chen, Xi, et al.
Published: (2025)
by: Chen, Xi, et al.
Published: (2025)
MERIT: Maximum-normalized Element-wise Ratio for Language Model Large-batch Training
by: Luo, Yang, et al.
Published: (2025)
by: Luo, Yang, et al.
Published: (2025)
Multimodal Multi-turn Conversation Stance Detection: A Challenge Dataset and Effective Model
by: Niu, Fuqiang, et al.
Published: (2024)
by: Niu, Fuqiang, et al.
Published: (2024)
How Does the Textual Information Affect the Retrieval of Multimodal In-Context Learning?
by: Luo, Yang, et al.
Published: (2024)
by: Luo, Yang, et al.
Published: (2024)
Bridge then Begin Anew: Generating Target-relevant Intermediate Model for Source-free Visual Emotion Adaptation
by: Zhu, Jiankun, et al.
Published: (2024)
by: Zhu, Jiankun, et al.
Published: (2024)
Open-Sora: Democratizing Efficient Video Production for All
by: Zheng, Zangwei, et al.
Published: (2024)
by: Zheng, Zangwei, et al.
Published: (2024)
MIPS at SemEval-2024 Task 3: Multimodal Emotion-Cause Pair Extraction in Conversations with Multimodal Language Models
by: Cheng, Zebang, et al.
Published: (2024)
by: Cheng, Zebang, et al.
Published: (2024)
Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning
by: Cheng, Zebang, et al.
Published: (2024)
by: Cheng, Zebang, et al.
Published: (2024)
OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models
by: Xue, Fuzhao, et al.
Published: (2024)
by: Xue, Fuzhao, et al.
Published: (2024)
A Closer Look at Time Steps is Worthy of Triple Speed-Up for Diffusion Model Training
by: Wang, Kai, et al.
Published: (2024)
by: Wang, Kai, et al.
Published: (2024)
Virtually Enriched NYU Depth V2 Dataset for Monocular Depth Estimation: Do We Need Artificial Augmentation?
by: Ignatov, Dmitry, et al.
Published: (2024)
by: Ignatov, Dmitry, et al.
Published: (2024)
AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models
by: Lian, Zheng, et al.
Published: (2025)
by: Lian, Zheng, et al.
Published: (2025)
Learning Transformer-based World Models with Contrastive Predictive Coding
by: Burchi, Maxime, et al.
Published: (2025)
by: Burchi, Maxime, et al.
Published: (2025)
MuDreamer: Learning Predictive World Models without Reconstruction
by: Burchi, Maxime, et al.
Published: (2024)
by: Burchi, Maxime, et al.
Published: (2024)
Higher fidelity perceptual image and video compression with a latent conditioned residual denoising diffusion model
by: Brenig, Jonas, et al.
Published: (2025)
by: Brenig, Jonas, et al.
Published: (2025)
Learned Lightweight Smartphone ISP with Unpaired Data
by: Arhire, Andrei, et al.
Published: (2025)
by: Arhire, Andrei, et al.
Published: (2025)
Practical Manipulation Model for Robust Deepfake Detection
by: Hopf, Benedikt, et al.
Published: (2025)
by: Hopf, Benedikt, et al.
Published: (2025)
Accurate and Efficient World Modeling with Masked Latent Transformers
by: Burchi, Maxime, et al.
Published: (2025)
by: Burchi, Maxime, et al.
Published: (2025)
Towards Online Real-Time Memory-based Video Inpainting Transformers
by: Thiry, Guillaume, et al.
Published: (2024)
by: Thiry, Guillaume, et al.
Published: (2024)
Neural Network Diffusion
by: Wang, Kai, et al.
Published: (2024)
by: Wang, Kai, et al.
Published: (2024)
Helen: Optimizing CTR Prediction Models with Frequency-wise Hessian Eigenvalue Regularization
by: Zhu, Zirui, et al.
Published: (2024)
by: Zhu, Zirui, et al.
Published: (2024)
SZTU-CMU at MER2024: Improving Emotion-LLaMA with Conv-Attention for Multimodal Emotion Recognition
by: Cheng, Zebang, et al.
Published: (2024)
by: Cheng, Zebang, et al.
Published: (2024)
Multi-source Domain Adaptation for Panoramic Semantic Segmentation
by: Jiang, Jing, et al.
Published: (2024)
by: Jiang, Jing, et al.
Published: (2024)
Dynamic Policy-Driven Adaptive Multi-Instance Learning for Whole Slide Image Classification
by: Zheng, Tingting, et al.
Published: (2024)
by: Zheng, Tingting, et al.
Published: (2024)
FaceChain-FACT: Face Adapter with Decoupled Training for Identity-preserved Personalization
by: Yu, Cheng, et al.
Published: (2024)
by: Yu, Cheng, et al.
Published: (2024)
DSP: Dynamic Sequence Parallelism for Multi-Dimensional Transformers
by: Zhao, Xuanlei, et al.
Published: (2024)
by: Zhao, Xuanlei, et al.
Published: (2024)
FaceChain-SuDe: Building Derived Class to Inherit Category Attributes for One-shot Subject-Driven Generation
by: Qiao, Pengchong, et al.
Published: (2024)
by: Qiao, Pengchong, et al.
Published: (2024)
MIORe & VAR-MIORe: Benchmarks to Push the Boundaries of Restoration
by: Ciubotariu, George, et al.
Published: (2025)
by: Ciubotariu, George, et al.
Published: (2025)
Cat: Post-Training Quantization Error Reduction via Cluster-based Affine Transformation
by: Zoljodi, Ali, et al.
Published: (2025)
by: Zoljodi, Ali, et al.
Published: (2025)
UAIC_Twin_Width: An Exact yet Efficient Twin-Width Algorithm
by: Arhire, Andrei, et al.
Published: (2025)
by: Arhire, Andrei, et al.
Published: (2025)
Resource-Efficient Iterative LLM-Based NAS with Feedback Memory
by: Gu, Xiaojie, et al.
Published: (2026)
by: Gu, Xiaojie, et al.
Published: (2026)
Babel-ImageNet: Massively Multilingual Evaluation of Vision-and-Language Representations
by: Geigle, Gregor, et al.
Published: (2023)
by: Geigle, Gregor, et al.
Published: (2023)
Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?
by: Geigle, Gregor, et al.
Published: (2024)
by: Geigle, Gregor, et al.
Published: (2024)
The Return of Structural Handwritten Mathematical Expression Recognition
by: Seitz, Jakob, et al.
Published: (2025)
by: Seitz, Jakob, et al.
Published: (2025)
LLM as a Neural Architect: Controlled Generation of Image Captioning Models Under Strict API Contracts
by: Jesani, Krunal, et al.
Published: (2025)
by: Jesani, Krunal, et al.
Published: (2025)
From Brute Force to Semantic Insight: Performance-Guided Data Transformation Design with LLMs
by: Shrestha, Usha, et al.
Published: (2026)
by: Shrestha, Usha, et al.
Published: (2026)
AugmentGest: Can Random Data Cropping Augmentation Boost Gesture Recognition Performance?
by: Aboudeshish, Nada, et al.
Published: (2025)
by: Aboudeshish, Nada, et al.
Published: (2025)
From Code to Prediction: Fine-Tuning LLMs for Neural Network Performance Classification in NNGPT
by: Hanouneh, Mahmoud, et al.
Published: (2026)
by: Hanouneh, Mahmoud, et al.
Published: (2026)
Similar Items
-
PEBench: A Fictitious Dataset to Benchmark Machine Unlearning for Multimodal Large Language Models
by: Xu, Zhaopan, et al.
Published: (2025) -
MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification
by: Xu, Zhaopan, et al.
Published: (2025) -
DAM: Dual Active Learning with Multimodal Foundation Model for Source-Free Domain Adaptation
by: Chen, Xi, et al.
Published: (2025) -
MERIT: Maximum-normalized Element-wise Ratio for Language Model Large-batch Training
by: Luo, Yang, et al.
Published: (2025) -
Multimodal Multi-turn Conversation Stance Detection: A Challenge Dataset and Effective Model
by: Niu, Fuqiang, et al.
Published: (2024)