Back to research

Research & Publications

Selected work on reliable generative models

Research on hallucination detection in language models and knowledge conflicts in diffusion models, alongside earlier work on automated exam evaluation.

The Detectability Gap: Hidden Heterogeneity in Hallucination Detection Across Language Models

GlobalSouthAI Workshop @ NeurIPS 2026

Accepted

Across four language models and three factual QA datasets, evaluated 12 model-dataset settings using lexical, semantic, and trajectory-level response dispersion. Identified high-agreement "Ghost" and low-agreement "Flickering" regimes, with a 0.35–0.46 AUC detectability gap.

  • Four language models and three factual QA datasets
  • 12 model-dataset settings evaluated
  • 0.35–0.46 AUC detectability gap

The Temporal Tug-of-War: Visualizing and Detecting RAG Conflicts in Diffusion Models via Trajectory Variance

UncertaiNLP Workshop @ EMNLP 2026

Accepted

Developed Trajectory Variance Score (TVS), measuring semantic divergence across independent diffusion trajectories. On LLaDA across four datasets, TVS achieved 70.10% accuracy and 0.7647 AUROC.

  • Trajectory Variance Score (TVS)
  • 70.10% accuracy on LLaDA across four datasets
  • 0.7647 AUROC

Eyes All Around: Design and Analysis of 360-Degree LiDAR Perception Using Equivariant Feature Learning in Unstructured Traffic

arXiv (2026)

arXiv

This paper studies a 360-degree LiDAR perception pipeline for autonomous driving, combining sector-wise panoramic processing with rotation-equivariant sparse convolutions. It evaluates a custom Ouster OS0 dataset collected across diverse Indian urban traffic conditions, finding strong detection for cars, buses, and trucks and lower scores for smaller, more variable road users.

Intellectual Property Rights and Entrepreneurship in the NFT Ecosystem: Legal Frameworks, Business Models, and Innovation Opportunities

arXiv (2025)

arXiv

This research examines the gap between traditional copyright law and blockchain-based NFT transactions. Using a mixed-methods approach, it introduces an IP rights matrix and a business-model taxonomy, and analyzes legal cases, smart contracts, and stakeholder interviews to identify challenges in cross-border enforcement, license standardization, and sustainable commercial opportunities.

Leveraging LLM and RAG for Automated Answer Script Evaluation

CSITSS 2024

IEEE Xplore

An Operating Systems answer-script evaluator using a fine-tuned LLaMA 2 model, handwriting recognition, and retrieval-augmented generation to ground grading in textbook content. Deployed with AWS SageMaker and Lambda; the project is described in the IEEE CSITSS 2024 publication.

Research interests

AI Safety Hallucination Detection Large Language Models Diffusion Language Models Mechanistic Interpretability Retrieval-Augmented Generation Reliable and Trustworthy AI