Project
Adaptive LLM Inference Engine
A resource-constrained system for choosing model size, quantization level, and execution settings based on available device capacity.
Backend AI, Built Right
Focused on Backend Development and AI/ML Engineering for tech enthusiasts and potential collaborators. Python, FastAPI, APIs, databases, LLM applications, and resource-aware AI systems shape the work.
Current focus
Python, FastAPI, APIs, databases, model integration, and deployment-friendly engineering.
Exploration
RAG, LangChain, LangGraph, prompt engineering, local models, quantization, and resource-aware inference.
Personal mission statement
Santosh M focuses on backend development and AI/ML engineering with a practical bias: shipping APIs, integrating models, and shaping tools that remain maintainable when the prototype ends.
FastAPI services, database-backed workflows, API design, and integration work that keeps product logic clean.
LLM applications, RAG, LangChain, LangGraph, AutoML, and resource-aware model selection for practical use cases.
Prompt design, evaluation, and pipeline thinking for AI features that feel useful rather than fragile.
Detailed project deep dives
A small set of projects that show the bridge between backend systems, AI engineering, and product thinking.

Academic project

Backend development

Product development

AI integrations

Learning project

Systems design
Skills and tools used
Santosh M works across the product surface and the service layer, with enough breadth to move from idea to implementation without losing sight of maintainability.
Backend
Service design, request handling, validation, and clean integration points.
Data
Data modeling, query-aware workflows, and retrieval support for AI use cases.
AI
Application patterns for search, generation, evaluation, and assistant flows.
Research
Model optimization, resource-aware execution, and deployment-sensitive tradeoffs.
Case studies of previous work
These examples focus on the shape of the work: what was being solved, how the stack was approached, and where AI fit into the system.
Project
A resource-constrained system for choosing model size, quantization level, and execution settings based on available device capacity.
Project
Backend development, AI integrations, prompt engineering, deployment plumbing, and the pipelines that support website generation.
Project
Exploration around RAG, LangGraph, local models, and the practical tradeoffs involved in building applications that stay responsive.
Technical blog articles
A small editorial feed for the ideas, experiments, and systems questions that shape the work.

How the routing, schema, and service boundaries come together in small but maintainable systems.

A practical view of model size, quantization, and execution constraints for edge and lightweight deployments.

How prompt design, feedback loops, and deployment decisions influence the stability of AI features.
About
Santosh M is building a direction around backend development and AI/ML engineering. The focus stays on practical product work: APIs, data flow, model integration, and the glue that lets tools be shipped and maintained.
Current learning spans LLM applications, RAG, LangChain, LangGraph, AutoML, generative AI, local models, quantization, and resource-aware inference. The goal is to contribute to teams that value technical clarity and thoughtful implementation.
Focus
A steady mix of implementation, experimentation, and learning — aimed at collaborators who want thoughtful execution rather than surface-level features.
Get in touch
This is the only direct action on the site. Use it to start a conversation about backend systems, AI features, or a shared build.
Newsletter sign-up
No extra capture funnels. This section exists as a light editorial placeholder for occasional updates, experiments, or write-ups.
Updates are occasional and focused on backend work, AI systems, and practical learning.