CABENCH: Benchmarking Composable AI for Solving Complex Tasks through Composing Ready-to-Use Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Pham, Tung-Thuy, Luong, Duy-Quan, Duong, Minh-Quan, Nguyen, Trung-Hieu, Nguyen, Thu-Trang, Nguyen, Son, Vo, Hieu Dinh |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Generating Critical Scenarios for Testing Automated Driving Systems
di: Nguyen, Trung-Hieu, et al.
Pubblicazione: (2024)
di: Nguyen, Trung-Hieu, et al.
Pubblicazione: (2024)
An Empirical Study on Capability of Large Language Models in Understanding Code Semantics
di: Nguyen, Thu-Trang, et al.
Pubblicazione: (2024)
di: Nguyen, Thu-Trang, et al.
Pubblicazione: (2024)
Model-Agnostic Correctness Assessment for LLM-Generated Code via Dynamic Internal Representation Selection
di: Vu, Thanh Trong, et al.
Pubblicazione: (2025)
di: Vu, Thanh Trong, et al.
Pubblicazione: (2025)
Reinforcement Learning-Based REST API Testing with Multi-Coverage
di: Nguyen, Tien-Quang, et al.
Pubblicazione: (2024)
di: Nguyen, Tien-Quang, et al.
Pubblicazione: (2024)
Correctness Assessment of Code Generated by Large Language Models Using Internal Representations
di: Bui, Tuan-Dung, et al.
Pubblicazione: (2025)
di: Bui, Tuan-Dung, et al.
Pubblicazione: (2025)
iML: Executable, Problem-Grounded, and Broadly Exploratory Code-Driven AutoML
di: Le, Dat, et al.
Pubblicazione: (2026)
di: Le, Dat, et al.
Pubblicazione: (2026)
Automated Description Generation for Software Patches
di: Vu, Thanh Trong, et al.
Pubblicazione: (2024)
di: Vu, Thanh Trong, et al.
Pubblicazione: (2024)
RAMBO: Enhancing RAG-based Repository-Level Method Body Completion
di: Bui, Tuan-Dung, et al.
Pubblicazione: (2024)
di: Bui, Tuan-Dung, et al.
Pubblicazione: (2024)
Finetuning LLMs for Automatic Form Interaction on Web-Browser in Selenium Testing Framework
di: Le, Nguyen-Khang, et al.
Pubblicazione: (2025)
di: Le, Nguyen-Khang, et al.
Pubblicazione: (2025)
Towards Test Generation from Task Description for Mobile Testing with Multi-modal Reasoning
di: Huynh, Hieu, et al.
Pubblicazione: (2025)
di: Huynh, Hieu, et al.
Pubblicazione: (2025)
Automated Web Application Testing: End-to-End Test Case Generation with Large Language Models and Screen Transition Graphs
di: Le, Nguyen-Khang, et al.
Pubblicazione: (2025)
di: Le, Nguyen-Khang, et al.
Pubblicazione: (2025)
Structured Exploration and Exploitation of Label Functions for Automated Data Annotation
di: Lam, Phong, et al.
Pubblicazione: (2026)
di: Lam, Phong, et al.
Pubblicazione: (2026)
Segment-Based Test Case Prioritization: A Multi-objective Approach
di: Huynh, Hieu, et al.
Pubblicazione: (2024)
di: Huynh, Hieu, et al.
Pubblicazione: (2024)
Impact of Code Transformation on Detection of Smart Contract Vulnerabilities
di: Manh, Cuong Tran, et al.
Pubblicazione: (2024)
di: Manh, Cuong Tran, et al.
Pubblicazione: (2024)
Early-Stage Prediction of Review Effort in AI-Generated Pull Requests
di: Minh, Dao Sy Duy, et al.
Pubblicazione: (2026)
di: Minh, Dao Sy Duy, et al.
Pubblicazione: (2026)
RBCTest: Leveraging LLMs to Mine and Verify Oracles of API Response Bodies for RESTful API Testing
di: Huynh, Hieu, et al.
Pubblicazione: (2025)
di: Huynh, Hieu, et al.
Pubblicazione: (2025)
Reconfigurable Intelligent Surfaces-assisted Positioning in Integrated Sensing and Communication Systems
di: Ta, Huyen-Trang, et al.
Pubblicazione: (2026)
di: Ta, Huyen-Trang, et al.
Pubblicazione: (2026)
Fourier-Attentive Representation Learning: A Fourier-Guided Framework for Few-Shot Generalization in Vision-Language Models
di: Pham, Hieu Dinh Trung, et al.
Pubblicazione: (2025)
di: Pham, Hieu Dinh Trung, et al.
Pubblicazione: (2025)
MEMRES: A Memory-Augmented Resolver with Confidence Cascade for Agentic Python Dependency Resolution
di: Minh, Dao Sy Duy, et al.
Pubblicazione: (2026)
di: Minh, Dao Sy Duy, et al.
Pubblicazione: (2026)
Toward Generation of Test Cases from Task Descriptions via History-aware Planning
di: Cao, Duy, et al.
Pubblicazione: (2025)
di: Cao, Duy, et al.
Pubblicazione: (2025)
ClozeMath: Improving Mathematical Reasoning in Language Models by Learning to Fill Equations
di: Pham, Quang Hieu, et al.
Pubblicazione: (2025)
di: Pham, Quang Hieu, et al.
Pubblicazione: (2025)
CodeLSI: Leveraging Foundation Models for Automated Code Generation with Low-Rank Optimization and Domain-Specific Instruction Tuning
di: Le, Huy, et al.
Pubblicazione: (2025)
di: Le, Huy, et al.
Pubblicazione: (2025)
An Approach of Structure‐Enhanced Code‐Centric Graph Learning for Just‐in‐Time Software Vulnerability Detection
di: Phu Pham, et al.
Pubblicazione: (2026)
di: Phu Pham, et al.
Pubblicazione: (2026)
Metric constructions and fixed point theorems in product spaces
di: Hieu, Doan Huu, et al.
Pubblicazione: (2026)
di: Hieu, Doan Huu, et al.
Pubblicazione: (2026)
Robust Aggregation for Federated Sequential Recommendation with Sparse and Poisoned Data
di: Nguyen, Minh Hieu
Pubblicazione: (2026)
di: Nguyen, Minh Hieu
Pubblicazione: (2026)
ViConsFormer: Constituting Meaningful Phrases of Scene Texts using Transformer-based Method in Vietnamese Text-based Visual Question Answering
di: Nguyen, Nghia Hieu, et al.
Pubblicazione: (2024)
di: Nguyen, Nghia Hieu, et al.
Pubblicazione: (2024)
Current status and trends of mangroves and seagrasses in Nha Trang bay
di: Nguyen, Xuan Hoa, et al.
Pubblicazione: (2015)
di: Nguyen, Xuan Hoa, et al.
Pubblicazione: (2015)
Task-driven Layerwise Additive Activation Intervention
di: Nguyen, Hieu Trung, et al.
Pubblicazione: (2025)
di: Nguyen, Hieu Trung, et al.
Pubblicazione: (2025)
Morphological variation and haplotype diversity of Halimeda macroloba and H. opuntia (Chlorophyta: Halimedaceae) from Southern Vietnam
di: Nguyen, Trung Hieu, et al.
Pubblicazione: (2022)
di: Nguyen, Trung Hieu, et al.
Pubblicazione: (2022)
W2E (Workout to Earn): A Low Cost DApp based on ERC-20 and ERC-721 standards
di: Son, Do Hai, et al.
Pubblicazione: (2024)
di: Son, Do Hai, et al.
Pubblicazione: (2024)
Cold-start Recommendation by Personalized Embedding Region Elicitation
di: Nguyen, Hieu Trung, et al.
Pubblicazione: (2024)
di: Nguyen, Hieu Trung, et al.
Pubblicazione: (2024)
Technique for Harvest of Superficial Quadriceps Tendon Autograft
di: Nguyen Quang Ton Quyen, et al.
Pubblicazione: (2024)
di: Nguyen Quang Ton Quyen, et al.
Pubblicazione: (2024)
VinDr-CXR-VQA: A Visual Question Answering Dataset for Explainable Chest X-Ray Analysis with Multi-Task Learning
di: Nguyen, Dang H., et al.
Pubblicazione: (2025)
di: Nguyen, Dang H., et al.
Pubblicazione: (2025)
MERVIN: A Unified Framework for Multimodal Event Retrieval in Vietnamese News Videos
di: Pham-Nguyen, Anh-Tai, et al.
Pubblicazione: (2026)
di: Pham-Nguyen, Anh-Tai, et al.
Pubblicazione: (2026)
SpareCodeSearch: Searching for Code Context When You Have No Spare GPU
di: Nguyen, Minh
Pubblicazione: (2025)
di: Nguyen, Minh
Pubblicazione: (2025)
Defect Prediction with Content-based Features
di: Pham, Hung Viet, et al.
Pubblicazione: (2024)
di: Pham, Hung Viet, et al.
Pubblicazione: (2024)
VN-MTEB: Vietnamese Massive Text Embedding Benchmark
di: Pham, Loc, et al.
Pubblicazione: (2025)
di: Pham, Loc, et al.
Pubblicazione: (2025)
Improving Vietnamese Legal Document Retrieval using Synthetic Data
di: Tien, Son Pham, et al.
Pubblicazione: (2024)
di: Tien, Son Pham, et al.
Pubblicazione: (2024)
Evaluating Classical Software Process Models as Coordination Mechanisms for LLM-Based Software Generation
di: Ha, Duc Minh, et al.
Pubblicazione: (2025)
di: Ha, Duc Minh, et al.
Pubblicazione: (2025)
Tracking Software Security Topics
di: Vu, Phong Minh, et al.
Pubblicazione: (2024)
di: Vu, Phong Minh, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Generating Critical Scenarios for Testing Automated Driving Systems
di: Nguyen, Trung-Hieu, et al.
Pubblicazione: (2024) -
An Empirical Study on Capability of Large Language Models in Understanding Code Semantics
di: Nguyen, Thu-Trang, et al.
Pubblicazione: (2024) -
Model-Agnostic Correctness Assessment for LLM-Generated Code via Dynamic Internal Representation Selection
di: Vu, Thanh Trong, et al.
Pubblicazione: (2025) -
Reinforcement Learning-Based REST API Testing with Multi-Coverage
di: Nguyen, Tien-Quang, et al.
Pubblicazione: (2024) -
Correctness Assessment of Code Generated by Large Language Models Using Internal Representations
di: Bui, Tuan-Dung, et al.
Pubblicazione: (2025)