Building RAG Pipelines for Local Llama
2026-01-15Quantized inference, chunking strategies, two-pass retrieval, and streaming chat UX for a fully local documentation assistant.
CSG, Government of Karnataka
Cyber Physical Systems Lab
FiXitAI
AgentStat
Real-time clinical note generation through a multi-stage STT to LLM reasoning pipeline. Built with PyTorch, Cerebras AI, and Llama 3.1 8B, including multi-agent orchestration and medical entity extraction.
Agentic, self-optimizing fact-checking pipeline with hybrid retrieval (FAISS + TF-IDF), claim extraction with spaCy/BART, and Streamlit-based verification interfaces.
Local PyTorch Llama inference with quantized models (GPTQ/4-bit), embedding-based search, streaming chat UI, and containerized deployment architecture for developer documentation.
Vision Transformer (ViT) with LoRA for 5-class diabetic retinopathy grading on MESSIDOR, achieving >80% validation accuracy. Includes attention rollout for lesion-level interpretability and failure mode analysis.
Quantized inference, chunking strategies, two-pass retrieval, and streaming chat UX for a fully local documentation assistant.
From CNN baselines to ViT+LoRA: handling class imbalance, parameter-efficient fine-tuning, and interpretability on diabetic retinopathy datasets.
Claim extraction with spaCy/BART, FAISS + TF-IDF hybrid retrieval, and self-optimizing verification loops that improve factual reliability.
Open to collaboration, consulting, and interesting engineering problems. Reach out directly or use the form.