Deesha AI · SkillLab
Break It Before Users Do — AI Reliability Lab
Two hands-on days making an AI application's behaviour measurable and dependable — an eval suite that catches regressions, retrieval and agent evaluation done separately, guardrails and injection defence, and the tracing that puts a cost and latency number on every answer.
Designed for
- Teams whose AI feature works in the demo and nobody can say how often
- Developers shipping RAG or agents with no regression net underneath
- Engineers asked "is it safe" and answering with an anecdote
- Anyone whose AI bill arrived as a surprise
- Fee
- ₹2,700
- Duration
- 2 days · 16 hours
- Each day
- 09:00 – 17:00
- Mode
- Offline / Online
The 2 days, hour by hour
10 hands-on sessions — every one ends in a thing
Day 1
09:00 – 17:00Measure it
5 sessions09:00 – 10:15
Find out how often it is actually wrong
why demos lie, eval datasets from real usage, baseline runs
Takeaway Your own application's real accuracy measured for the first time — a number, not an impression, and usually a humbling one.
10:30 – 11:45
Score answers without a human reading each one
exact checks vs model graders, rubric design, when graders lie
Takeaway An automatic scorer that agrees with your own judgment — validated against answers you graded by hand first.
12:00 – 13:00
Evaluate retrieval before blaming the model
retrieval metrics, golden documents, fixing the right layer
Takeaway A wrong answer diagnosed to retrieval, not generation — proved by the fix landing in the retriever and the score moving.
14:00 – 15:15
Evaluate what an agent did, not just what it said
trajectory evaluation, tool-choice correctness, stopping behaviour
Takeaway An agent scored on its decisions — right tool, right order, knew when to stop — across twenty recorded runs.
15:30 – 17:00
Build the failure catalogue
failure taxonomy, adversarial cases, the regression set
Takeaway Your system broken deliberately in ten distinct ways — each failure named, captured and added to the regression set.
Day 2
09:00 – 17:00Control it
5 sessions09:00 – 10:15
Guard the inputs
input validation, prompt injection, scope refusal
Takeaway An injection attempt that worked yesterday caught at the gate today — and the off-topic request politely refused.
10:30 – 11:45
Guard the outputs
output schemas, unsafe content checks, grounding enforcement
Takeaway An ungrounded claim stopped before the user saw it — the answer forced back to evidence or withheld.
12:00 – 13:00
Trace every request end to end
tracing spans, token and cost attribution, latency budgets
Takeaway One slow, expensive conversation dissected span by span — the cost attributed to the exact call that earned it.
14:00 – 15:15
Put cost and latency on a budget
cost ceilings, latency targets, degrading deliberately
Takeaway A runaway conversation hitting a ceiling you set — and degrading to a cheaper path instead of a bigger bill.
15:30 – 17:00
Gate the pipeline and defend the numbers
evals in CI, release criteria, the reliability review
Takeaway A deliberate quality regression blocked by your own CI gate — and your system's reliability defended to the room in numbers.
By the end of day 2, you are holding
A reliability harness around your own AI application — an eval dataset and scoring that run in CI, retrieval and agent behaviour measured separately, guardrails that catch an injection attempt live, and a trace dashboard that prices every conversation — plus the failure catalogue you built by breaking it yourself.
Request a schedule
Break It Before Users Do — AI Reliability Lab
Runs on request, for individuals and for teams.
More SkillLabs
One weekend, one capability
Agent Builder Lab
₹2,700 · 2 days · 16 hours
Runs on request
See the plan
AI App Engineering Lab
₹2,700 · 2 days · 16 hours
Runs on request
See the plan
Ground the Model — RAG Engineering Lab
₹2,700 · 2 days · 16 hours
Runs on request
See the plan
Agents in the Wild — Production Agent Lab
₹2,700 · 2 days · 16 hours
Runs on request
See the plan
Give Your Agent Hands — MCP Tooling Lab
₹2,700 · 2 days · 16 hours
Runs on request
See the plan
Docker & Kubernetes for Developers
₹2,700 · 2 days · 16 hours
Next: 19 Sept
See the plan
Git, GitHub and GitHub Actions for Developers
₹2,700 · 2 days · 16 hours
Next: 22 Aug
See the plan
Ship It on AWS — Cloud Deployment for Developers
₹2,700 · 2 days · 16 hours
Runs on request
See the plan
SQL for Application Developers
₹2,700 · 2 days · 16 hours
Next: 29 Aug
See the plan
Events in Motion — Kafka & Messaging for Java Developers
₹2,700 · 2 days · 16 hours
Runs on request
See the plan
Linux Command Line Essentials
₹2,700 · 2 days · 16 hours
Runs on request
See the plan
Logging & Monitoring for Production
₹2,700 · 2 days · 16 hours
Runs on request
See the plan
Performance & Production Readiness
₹2,700 · 2 days · 16 hours
Runs on request
See the plan
Redis & Caching for Java Services
₹2,700 · 2 days · 16 hours
Runs on request
See the plan
Secure the Frontend — Authentication & Web Security Lab
₹2,700 · 2 days · 16 hours
Runs on request
See the plan
Test It Before You Ship It — Java API Testing Lab
₹2,700 · 2 days · 16 hours
Runs on request
See the plan
Test the User Journey — Frontend Testing Lab
₹2,700 · 2 days · 16 hours
Runs on request
See the plan
TypeScript for Application Developers
₹2,700 · 2 days · 16 hours
Runs on request
See the plan
Who runs it
Practitioners, in the room with you
All trainers and mentorsEach workshop names its trainer before you book.

Vishal Shah
Founder & Principal Trainer
Two decades building and teaching commerce, banking and cloud platforms — still writing code
- Java & Spring Boot microservices
- TypeScript, React, Angular & Next.js
- Composable commerce (commercetools)

Shrenik Shah
Principal Trainer
Cloud and AI architect who has upskilled over 5,000 engineers in Java, React, Python and cloud
- Python — Django, Flask, GenAI & agentic workflows
- Java & Spring Boot
- React & Angular, micro-frontend architecture
Running this for a team?
We deliver SkillLabs on site for institutions and engineering teams.
₹2,700
2 days · 16 hours