Hey everyone!
AI / Full-Stack Engineer focused on shipping production-grade LLM + RAG systems
### What I build:
• Reliable RAG pipelines for accurate, grounded search & Q&A
• Automated summarization + structured extraction
• Domain-specific chat agents that actually work in production
• Full end-to-end LLM flows (ingestion → retrieval → generation → deployment)
### Stack I usually work with:
• Python / TypeScript • FastAPI / Express • SSE streaming
• OpenAI, Anthropic, Gemini, Llama 3/3.1, Mistral family
• LangChain / LlamaIndex • chunking strategies, hybrid search, re-ranking
• Vector DBs: Pinecone, Qdrant, FAISS, OpenSearch
• Infra: Docker/K8s, vLLM, AWS/GCP, proper CI/CD
If you're working on RAG, agents, fine-tuning, or scaling LLMs in production and want to bounce ideas / collab / need help debugging something — feel free to ping me. Happy to share war stories or look at your setup
Looking forward to chatting!