Deesha Tech Academy
Menu

Deesha AI · SkillLab

WorkshopRuns on requestAI & Agentic AIEvalsGuardrailsTracingRAG evaluationAgent evaluation

Break It Before Users Do — AI Reliability Lab

Two hands-on days making an AI application's behaviour measurable and dependable — an eval suite that catches regressions, retrieval and agent evaluation done separately, guardrails and injection defence, and the tracing that puts a cost and latency number on every answer.

Designed for

  • Teams whose AI feature works in the demo and nobody can say how often
  • Developers shipping RAG or agents with no regression net underneath
  • Engineers asked "is it safe" and answering with an anecdote
  • Anyone whose AI bill arrived as a surprise
Fee
₹2,700
Duration
2 days · 16 hours
Each day
09:00 – 17:00
Mode
Offline / Online
Request a schedule

The 2 days, hour by hour

10 hands-on sessions — every one ends in a thing

Day 1

09:00 – 17:00

Measure it

5 sessions

09:00 – 10:15

Find out how often it is actually wrong

why demos lie, eval datasets from real usage, baseline runs

Takeaway Your own application's real accuracy measured for the first time — a number, not an impression, and usually a humbling one.

10:30 – 11:45

Score answers without a human reading each one

exact checks vs model graders, rubric design, when graders lie

Takeaway An automatic scorer that agrees with your own judgment — validated against answers you graded by hand first.

12:00 – 13:00

Evaluate retrieval before blaming the model

retrieval metrics, golden documents, fixing the right layer

Takeaway A wrong answer diagnosed to retrieval, not generation — proved by the fix landing in the retriever and the score moving.

14:00 – 15:15

Evaluate what an agent did, not just what it said

trajectory evaluation, tool-choice correctness, stopping behaviour

Takeaway An agent scored on its decisions — right tool, right order, knew when to stop — across twenty recorded runs.

15:30 – 17:00

Build the failure catalogue

failure taxonomy, adversarial cases, the regression set

Takeaway Your system broken deliberately in ten distinct ways — each failure named, captured and added to the regression set.

Day 2

09:00 – 17:00

Control it

5 sessions

09:00 – 10:15

Guard the inputs

input validation, prompt injection, scope refusal

Takeaway An injection attempt that worked yesterday caught at the gate today — and the off-topic request politely refused.

10:30 – 11:45

Guard the outputs

output schemas, unsafe content checks, grounding enforcement

Takeaway An ungrounded claim stopped before the user saw it — the answer forced back to evidence or withheld.

12:00 – 13:00

Trace every request end to end

tracing spans, token and cost attribution, latency budgets

Takeaway One slow, expensive conversation dissected span by span — the cost attributed to the exact call that earned it.

14:00 – 15:15

Put cost and latency on a budget

cost ceilings, latency targets, degrading deliberately

Takeaway A runaway conversation hitting a ceiling you set — and degrading to a cheaper path instead of a bigger bill.

15:30 – 17:00

Gate the pipeline and defend the numbers

evals in CI, release criteria, the reliability review

Takeaway A deliberate quality regression blocked by your own CI gate — and your system's reliability defended to the room in numbers.

By the end of day 2, you are holding

A reliability harness around your own AI application — an eval dataset and scoring that run in CI, retrieval and agent behaviour measured separately, guardrails that catch an injection attempt live, and a trace dashboard that prices every conversation — plus the failure catalogue you built by breaking it yourself.

Request a schedule

Break It Before Users Do — AI Reliability Lab

Runs on request, for individuals and for teams.

How will you join?

Laptop ready with the prerequisites?

Opens WhatsApp with a message naming this workshop. We confirm your seat and payment by reply — this website stores nothing.

More SkillLabs

One weekend, one capability

Browse the SkillLabs catalogue

Agent Builder Lab

₹2,700 · 2 days · 16 hours

Runs on request

See the plan

AI App Engineering Lab

₹2,700 · 2 days · 16 hours

Runs on request

See the plan

Ground the Model — RAG Engineering Lab

₹2,700 · 2 days · 16 hours

Runs on request

See the plan

Agents in the Wild — Production Agent Lab

₹2,700 · 2 days · 16 hours

Runs on request

See the plan

Give Your Agent Hands — MCP Tooling Lab

₹2,700 · 2 days · 16 hours

Runs on request

See the plan

Docker & Kubernetes for Developers

₹2,700 · 2 days · 16 hours

Next: 19 Sept

See the plan

Git, GitHub and GitHub Actions for Developers

₹2,700 · 2 days · 16 hours

Next: 22 Aug

See the plan

Ship It on AWS — Cloud Deployment for Developers

₹2,700 · 2 days · 16 hours

Runs on request

See the plan

SQL for Application Developers

₹2,700 · 2 days · 16 hours

Next: 29 Aug

See the plan

Events in Motion — Kafka & Messaging for Java Developers

₹2,700 · 2 days · 16 hours

Runs on request

See the plan

Linux Command Line Essentials

₹2,700 · 2 days · 16 hours

Runs on request

See the plan

Logging & Monitoring for Production

₹2,700 · 2 days · 16 hours

Runs on request

See the plan

Performance & Production Readiness

₹2,700 · 2 days · 16 hours

Runs on request

See the plan

Redis & Caching for Java Services

₹2,700 · 2 days · 16 hours

Runs on request

See the plan

Secure the Frontend — Authentication & Web Security Lab

₹2,700 · 2 days · 16 hours

Runs on request

See the plan

Test It Before You Ship It — Java API Testing Lab

₹2,700 · 2 days · 16 hours

Runs on request

See the plan

Test the User Journey — Frontend Testing Lab

₹2,700 · 2 days · 16 hours

Runs on request

See the plan

TypeScript for Application Developers

₹2,700 · 2 days · 16 hours

Runs on request

See the plan

Who runs it

Practitioners, in the room with you

All trainers and mentors

Each workshop names its trainer before you book.

  • Vishal Shah

    Vishal Shah

    Founder & Principal Trainer

    Two decades building and teaching commerce, banking and cloud platforms — still writing code

    • Java & Spring Boot microservices
    • TypeScript, React, Angular & Next.js
    • Composable commerce (commercetools)
    Full profile →
  • Shrenik Shah

    Shrenik Shah

    Principal Trainer

    Cloud and AI architect who has upskilled over 5,000 engineers in Java, React, Python and cloud

    • Python — Django, Flask, GenAI & agentic workflows
    • Java & Spring Boot
    • React & Angular, micro-frontend architecture
    Full profile →

Running this for a team?

We deliver SkillLabs on site for institutions and engineering teams.

Talk to Deesha

₹2,700

2 days · 16 hours

Request a schedule