About
Hi, Iām Tanmay Dipak Patil ā a Machine Learning Engineer with a focus on GPU programming, inference optimization, and diffusion research.
I currently work at ModelsLab, where I build and optimize inference interfaces for language, image, and audio models. I spend a lot of my time chasing lower latency, reading papers, and implementing ideas from scratch to go deep on diffusion research.
What I do
- ML / AI: TensorFlow, PyTorch, Transformers, Diffusers, Computer Vision
- GPU Programming: Triton, Cute-dsl, Hip, Hipkittens, CUDA
- Programming: Python, JavaScript, Dart, C++
- Tools: Git, Docker, Kubernetes, Django, FastAPI
Experience
ModelsLab ā Machine Learning Engineer (Feb 2024 ā Present)
- Developed inference interfaces for language, image, and audio models, plus services such as realtime chat, voice cloning, and image/video synthesis & editing. Benchmarked approaches for arbitrary model serving in real time and made existing implementations much faster to reduce generation latency.
- Trained and finetuned image generation models using LoRA, DPO, and RLHF, dedicating significant time to research and implementing ideas from scratch.
- Managed GPU deployments and handled multiple major production outages.
PandasAI ā Software Engineer Intern (Sep 2023 ā Jan 2024)
- Spearheaded development of open-source online connectors and streamlined pipeline construction for LLMs, enhancing data accessibility and deployment efficiency.
Intersense Technologies LLP ā Python Developer Intern (Feb 2022 ā Nov 2022)
- Built an ML-powered automated offset correction unit for CNC machines on a Raspberry Pi 4, replacing manual offset correction using Python, PyQt5, and network programming.
Projects
Krea 2 Depth ControlNet LoRA ā depth-conditioned image generation
- Developed and published a depth-conditioned ControlNet-LoRA for Krea-2, enabling image-to-image generation that preserves 3D structure and composition from input depth maps.
- Implemented latent-space depth control using Depth-Anything-V2, Qwen-Image VAE conditioning, and rank-64 LoRA adaptation, achieving 0.98ā0.99 depth consistency across generations.
Enigma DSL ā an MLIR-based GPU kernel compiler
- Built a Python GPU kernel DSL that traces into a custom MLIR dialect and lowers to native GPU shading language, applying CuTe-style layout algebra (composition, coalesce, TV layouts) to map threads/wavefronts to memory.
- Published to PyPI as
enigma-dsl.
ModelQ ā a lightweight, production-ready async task library
- Simplifies development and execution of asynchronous tasks in distributed systems. Inspired by Celery, it provides a clear API for defining, scheduling, and managing background jobs and complex task workflows ā handling millions of requests daily in production.
Education
CSMSS CSCOE ā BTech in AI and Data Science, CGPA 8.84/10.0 (July 2021 ā June 2024)
Achievements
- Open source contributor to trending ML repos including Hugging Face Transformers, Hipkittens, and PandasAI.
Get in touch
- š§ tanmaypatil3151@gmail.com
- š» GitHub
- š¤ Hugging Face
- š¦ Twitter/X
- š Resume