Research Output

Publications &
Ongoing Work

Research at the intersection of AI systems, software security, and machine learning — targeting high-impact venues and real-world applicability.

JOSS Submitted · February 2026

VibeBench: A Holistic Benchmarking Framework for Large Language Model Evaluation

Muktadir Arif et al. · Journal of Open Source Software

An open-source framework for holistic LLM evaluation spanning reasoning, instruction-following, creativity, and tool-use tasks. VibeBench includes curated multi-domain datasets, a multi-model evaluation pipeline, and a live leaderboard. Released as v1.0.0 under MIT license.

Python LLM Evaluation Benchmarking Open Source · MIT
In Progress Active · 2026

ML-Driven Spending Prediction for Hostel Environments

Muktadir Arif · Independent Research

Leveraging behavioral spending data collected from 70+ students via the Hostel Expense Tracker to train regression and time-series models for monthly expense prediction. Goal: deploy as an in-app module to provide personalized budgeting forecasts.

Scikit-Learn Python Firebase Time-Series

Status: Data collection & feature engineering phase