Harsh Gurawaliya
AI Engineer
Building production-ready AI systems: from multi-agent AI and RAG pipelines to computer-use agents.
Featured Whitepaper
Agentic Memory Realized
First-authored by me at appliedAI, this whitepaper is a comparative study of memory architectures for long-term AI agents, exploring how systems such as Zep, Letta, and Mem0 move autonomous agents from stateless interactions toward cognitive continuity.
About Me
I am a curious learner drawn to mathematics, statistics, and difficult problem solving, now using Python to build practical AI systems with agentic workflows. My current work focuses on AI agents, LLM memory, computer-use automation, RAG, and production-ready AI systems for real client use cases.
Alongside my role at appliedAI Initiative GmbH, I am completing a B.Sc. in Artificial Intelligence at Technische Hochschule Deggendorf. My bachelor thesis explores agentic memory and benchmarks different agentic AI frameworks, connecting my academic work directly with the systems I build professionally.
Experience
- First author of the white paper Agentic Memory Realized, comparing Zep, Letta, and Mem0 for long-horizon memory in autonomous AI agents.
- Built a computer-use agent powered by Claude for autonomous browser and desktop task execution in internal automation use cases.
- Designed and delivered a client-facing email agent using AWS Strands Agents and Amazon SES to automate inbound triage and outbound communication workflows.
- Worked across tool use, planning, memory, multi-agent orchestration, and evaluation for real client deliverables.
- Engineered a production-grade RAG application from scratch with FastAPI, LlamaIndex, Ollama, custom chunking, and ChromaDB retrieval.
- Improved document generation accuracy by 30% compared with standard ChatGPT on raw company documents while preserving privacy through local models.
- Architected backend systems with PostgreSQL, AWS S3, JWT-secured REST APIs, role-based access control, and OpenVPN-based local LLM infrastructure.
- Developed a post-processing method for deep neural network predictions focused on label error detection in semantic segmentation tasks.
- Implemented a YOLOv8 and PyTorch object detection pipeline; earned a Grade 1 in the Computer Vision course.
Education
Selected Projects
Built an agentic VC investment analyst that turns a one-line company description into an evidence-cited, source-traceable investment memo through a structured seven-stage dialectic: decompose, multi-source evidence, pro/con generation, adversarial critique, refine, quantitative scoring, and recommendation.
Engineered against LLM overconfidence with a mandatory devil's-advocate critique that stress-tests every argument, then rolls signals into a 0-100 composite score with an explicit confidence level.
Designed pluggable, declarative strategy profiles for dimensions, signal weights, rubrics, thresholds, connectors, and guardrails, allowing the same deal to be scored against multiple fund theses using signals from news, GitHub, hiring, web traffic, and funding sources.
Founded and built a patient-data platform that records doctor-patient consultations, transcribes them in real time with Whisper, and translates Tamil and Hindi into English.
Generated LLM-driven medical summaries from reports, transcripts, and lab data, with trend detection, similar-patient diagnosis assistance, role-based authentication, doctor review, follow-up scheduling, and a patient waiting-list queue.
Validated the product with six senior doctors in India, receiving positive testimonials on its potential to improve hospital workflows and patient outcomes.
Fine-tuned LLaMA 3 with QLoRA on a large patient-doctor conversation dataset for medical Q&A and triage scenarios.
Merged adapters into the base model, quantized it to GGUF, and deployed it locally in LM Studio for secure offline domain-specific conversational AI.
Trained a UNet with dropout and skip connections on 2,146 brain MRI scans using a custom Dice and Focal loss strategy to handle class imbalance and improve boundary precision.
Achieved 98% global accuracy and 98% mean IoU on the held-out test set.
Trained a PPO agent to land an orbital-class rocket in simulation, shaping rewards for stable touchdowns and tuning the policy network for consistent autonomous landings.
Built a complete MLOps loop including data ingestion, preprocessing, Scikit-learn training, a Flask and React app, Docker, GitHub Actions CI/CD, and AWS deployment for automated, versioned inference.
Skills
Contact Me
I'm open to discussing AI agents, LLM memory, RAG systems, healthcare AI products, and practical automation projects.