Undergraduate Researcher: LLM Evaluation & Benchmarking
- Curated and annotated high-quality, multi-domain benchmark datasets across diverse categories, including Medical and Religious domains, ensuring robust data preparation for model ingestion.
- Conducted comprehensive evaluations of state-of-the-art generative language models (e.g., LLaMA, Qwen, Mistral) to systematically track response accuracy and answer correctness.
- Analyzed model performance limitations and behavioral edge cases, focusing on measuring hallucination rates and evaluating domain-specific knowledge alignment.