The Detectability Gap: Hidden Heterogeneity in Hallucination Detection Across Language Models
GlobalSouthAI Workshop @ NeurIPS 2026
Across four language models and three factual QA datasets, evaluated 12 model-dataset settings using lexical, semantic, and trajectory-level response dispersion. Identified high-agreement "Ghost" and low-agreement "Flickering" regimes, with a 0.35–0.46 AUC detectability gap.
- Four language models and three factual QA datasets
- 12 model-dataset settings evaluated
- 0.35–0.46 AUC detectability gap
The Temporal Tug-of-War: Visualizing and Detecting RAG Conflicts in Diffusion Models via Trajectory Variance
UncertaiNLP Workshop @ EMNLP 2026
Developed Trajectory Variance Score (TVS), measuring semantic divergence across independent diffusion trajectories. On LLaDA across four datasets, TVS achieved 70.10% accuracy and 0.7647 AUROC.
- Trajectory Variance Score (TVS)
- 70.10% accuracy on LLaDA across four datasets
- 0.7647 AUROC
Eyes All Around: Design and Analysis of 360-Degree LiDAR Perception Using Equivariant Feature Learning in Unstructured Traffic
arXiv (2026)
This paper studies a 360-degree LiDAR perception pipeline for autonomous driving, combining sector-wise panoramic processing with rotation-equivariant sparse convolutions. It evaluates a custom Ouster OS0 dataset collected across diverse Indian urban traffic conditions, finding strong detection for cars, buses, and trucks and lower scores for smaller, more variable road users.
Intellectual Property Rights and Entrepreneurship in the NFT Ecosystem: Legal Frameworks, Business Models, and Innovation Opportunities
arXiv (2025)
This research examines the gap between traditional copyright law and blockchain-based NFT transactions. Using a mixed-methods approach, it introduces an IP rights matrix and a business-model taxonomy, and analyzes legal cases, smart contracts, and stakeholder interviews to identify challenges in cross-border enforcement, license standardization, and sustainable commercial opportunities.
Leveraging LLM and RAG for Automated Answer Script Evaluation
CSITSS 2024
An Operating Systems answer-script evaluator using a fine-tuned LLaMA 2 model, handwriting recognition, and retrieval-augmented generation to ground grading in textbook content. Deployed with AWS SageMaker and Lambda; the project is described in the IEEE CSITSS 2024 publication.