Skip to content

How-to Guides

Task guides for building an AI PDLC: a replication playbook based on the flagship disclosed implementation (Freshworks), recipes for adopting the off-the-shelf tools (GitHub Spec Kit, AWS Kiro), and the research and verification recipes used to maintain this topic folder.

Provenance Strata

The sections of this note deliberately mix strata. Each is labeled. [F] = Freshworks-reported fact from the flagship case study's primary source. [P] = publicly reproducible mechanism (GitHub Spec Kit / Kiro docs). [R] = replication guidance synthesized from disclosures using real, named tooling (generic-by-design. Freshworks' internal implementation remains unpublished). [V] = vault-side verification tooling for auditing this note set over time.

The Freshworks operating model (actors and their roles) is described in Explanation. This page covers the tasks.

Replication Playbook

A sequenced adoption path honoring the core ordering claim of the disclosure — foundation before harness:

  1. Encode the design system for agents — publish design tokens/components in machine-readable form (token JSON + component metadata). Acceptance test: an LLM asked to enumerate components misses none. The disclosed failure mode to guard: Figma Make silently skipped design-system components until humans caught it.
  2. Codify standards as ingestible rules — convert tribal engineering norms into written rule files your harness auto-loads (Cursor supports repo-level rules directories consumed at runtime. Equivalent concepts exist in every agentic editor).
  3. Consolidate the monorepo — one repository as sole source of truth so every phase reads/writes the same reality.
  4. Stand up the three Prism analogues — start thin: (a) knowledge index of products + dependencies, (b) per-feature state that threads between phases, (c) one shared directory of skills/rules/commands/agents. Thin versions beat no versions. Expand when phases leak context.
  5. Wire the entrypoint — a single slash command collecting business unit, Epic ID, and feature team, then walking phases idea-brief → prototype → QA [F-shape].
  6. Force requirements interrogation — the demoed behavior worth copying verbatim: the agent asks persona, drill-down depth, and success-criteria questions before generating artifacts.
  7. Ground PRDs in the warehouse — give agents a scoped, read-only analytics role. On Databricks (Unity Catalog), least-privilege starts at three separate grants: GRANT USE CATALOG ON CATALOG main TO \ai-agent-ro`;thenGRANT USE SCHEMA ON SCHEMA main.analytics TO `ai-agent-ro`;thenGRANT SELECT ON TABLE main.analytics.events TO `ai-agent-ro`;`. Every number shown in the generated PRD must trace to a query in the document, as demonstrated on Bel across 4,358 active ITSM accounts.
  8. Name a review gate after its owner — institutionalize executive-review-as-pipeline-step (theirs is literally called "CPO check"). Keep gate criteria written down so succession does not break the mechanism (see architecture risks).
  9. Terminate in evals — adopt a real eval framework as the last phase before release, for example promptfoo: npx promptfoo@latest init && npx promptfoo@latest eval against assertion suites per artifact type.
  10. Keep the harness model-agnostic — pin zero vendors. Benchmark models per-phase (theirs flipped Claude → Grok on observed speed alone).

Productize Instead Of Building From Zero [P/R]

Steps 4-9 exist off the shelf in Spec Kit (1.0.11, 2026-09-24). Adopt its chain before writing custom harness code. Requirements: Python 3.11+ and uv.

uv tool install specify-cli
specify init my-project --integration claude   # or copilot, codex, cursor-agent, gemini, kiro-cli ...
cd my-project
specify integration list                        # installed and available integrations

Then run the lifecycle inside the agent's chat, not in the terminal. The dotted names below are the reference form; skills-based agents such as Copilot and Claude Code use /speckit-constitution, Codex uses $speckit-constitution (invocation forms):

  1. /speckit.constitution once per project (principles, maps to Prism's written standards).
  2. /speckit.specify, then /speckit.clarify (the interrogation gate, max five questions per pass).
  3. /speckit.plan, then /speckit.checklist (reviewer-owned checklists: the human gate).
  4. /speckit.tasks, then /speckit.analyze (read-only consistency check; fix issues at the step that owns them).
  5. /speckit.implement, then /speckit.converge. Repeat until converge reports Converged.

The gap versus the flagship case that remains yours to build: warehouse-grounded quantitative evidence inside specs (step 7 has no Spec Kit analogue today).

Add Idea Assessment And Bug-Fix Processes [P]

Two opt-in first-party extensions cover the phases before and after feature delivery:

specify extension add assess   # intake -> research -> define -> shape -> decide
specify extension add bug      # assess -> fix -> test

In the agent: /speckit-assess-intake "Let users work offline." slug=offline-mode, then -research, -define, -shape, -decide with the same slug. Artifacts land in .specify/assessments/<slug>/ and end in a go / needs-clarification / kill decision. A go can be handed to /speckit-specify. Bug reports land in .specify/bugs/<slug>/ with a verified, partial or failed verdict.

Keep Spec Kit Current [P]

specify self check              # read-only: are you on the latest release?
specify self upgrade --dry-run  # show what would run
specify self upgrade            # upgrade in place (uv tool or pipx installs)
specify self upgrade --tag v1.0.11

Releases ship every few days, so pin a tag in CI and upgrade deliberately.

Write Kiro-Style EARS Acceptance Criteria [P/R]

Use EARS even outside Kiro: it makes requirements something an agent (and a property-based test) can check. For each user story in requirements.md:

  1. Write one criterion per observable behaviour, using a pattern from the EARS table. Most criteria are event-driven: WHEN <trigger> THE SYSTEM SHALL <response>.
  2. Add at least one unwanted-behaviour criterion per story: IF <failure> THEN THE SYSTEM SHALL <response>.
  3. Keep numbers explicit ("drops 20% week-over-week", not "drops a lot"), so the criterion becomes a test.
  4. Pick the Kiro variant that fits: a Requirements-First feature spec when behaviour is known, Design-First when architecture constraints drive the work, a Bugfix spec to record root cause and the behaviour that must not change.
  5. Review and approve requirements.md before design.md is generated. Skip the gates (Quick spec) only for low-risk work.

Commands And Recipes

Research provenance maintenance [V]

# Re-pull the primary source into working memory
defuddle parse https://www.news.aakashg.com/p/srini-raghavan-podcast --md -o /tmp/podcast-article.md

# Confirm episode video identity (title check) without loading YouTube's JS app
curl -s "https://www.youtube.com/oembed?url=https%3A//www.youtube.com/watch%3Fv%3D0qNZVlW8IR4&format=json" | python3 -c "import sys,json; print(json.load(sys.stdin)['title'])"

# Re-verify headline financial claim against Freshworks' own releases
curl -sL -A "Mozilla/5.0" https://www.freshworks.com/pressrelease/freshworks-reports-first-quarter-2026-results/ \
  | grep -ioE 'revenue[^<]{0,120}' | head -5

Hard-earned quirk recorded here: several CDN-fronted pages (Freshworks IR, TipRanks, GlobeNewswire, StockTitan) answer HEAD requests with 403/connection reset while serving browsers fine — verifying with curl -I produces false negatives.

# Liveness check that matches real traffic: full GET + browser UA
for u in $(grep -ohE 'https://[^ )>"|]+' *.md | sort -u); do
  code=$(curl -s -o /dev/null -w "%{http_code}" --max-time 12 -L \
    -A "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 Chrome/126.0 Safari/537.36" "$u")
  echo "$code $u"
done

Vault content-quality gates [V]

Note: never paste the banned-pattern literals of this vault into notes — the audits grep knowledge/ verbatim, so quoting them creates permanent false positives.

# Threshold sanity for the standard topic shape
# (index >= 80, explanation >= 150, how-to-guides >= 100, reference >= 80)
wc -l index.md explanation.md how-to-guides.md reference.md

# Anti-template sweep. Canonical ban-list lives in:
#   agents.md -> Content Quality Rules
#   .agents/skills/vault-maintenance/SKILL.md -> audit checklist item 10
# Expect zero hits; a hit means research-backed rewriting is required.
git grep -niE '(template filler|placeholder table)' -- knowledge/ai-agents/ai-pdlc/

# Child-title convention: plain page-type titles for the MkDocs sidebar
for f in explanation.md reference.md how-to-guides.md; do head -2 "$f" | tail -1; done

Artifact-library health checks [R]

No Freshworks-internal tooling is published for this. The check itself is generic and safe to run on any harness repo:

# Stale-artifact sweep: artifacts untouched while code moved underneath them
git log --format="%ct %H" -1 -- ./agents/artifacts/skill-perf-dashboard.md   # compare ages vs. src/

# Find artifacts referencing dead file paths (breakage candidates)
grep -rhoE '"(src|packages)/[^"]+"' ./agents/artifacts/*.md \
  | tr -d '"' | sort -u | while read p; do [ -e "$p" ] || echo "STALE REF: $p"; done

Known Issues And Failure Modes

Symptom Disclosed? First Response
Prototype generator skips design-system components Yes — happened live in Figma Make Component-enumeration acceptance test on every design-system change. Route gaps to humans (their stated position: "those are the places where humans still have a role")
Harness latency balloons on frontier models Observed contrast (Claude vs Grok, ≤10–15 s steps on Grok) Benchmark per phase. Reserve heavyweight models for evals/analysis passes rather than interactive loops
PRD drifts from quantitative reality Not observed publicly. Structurally possible Enforce query-traceability in generated documents. Sample-audit numbers monthly [R]
Spec Kit analyze reports gaps between spec, plan and tasks Documented behaviour Fix at the owning step (specify/clarify for requirements, plan for design, tasks to regenerate), then re-run analyze until clean [P]
Implementation stops to ask about unchecked checklist items Documented behaviour since checklists became reviewer-owned A reviewer ticks each item after checking the requirement. The agent must not tick them itself [P]
/speckit.* command not found in the agent Documented The spelling depends on the integration: /speckit-specify (skills mode), $speckit-specify (Codex), /skill:speckit-specify (Kimi) [P]
Gate quality decays when the person it is named after leaves Structurally realized — Raghavan's departure was announced 2026-07-28, after the disclosure was recorded Rewritten gate criteria owned by role, not individual. See the threat model

Cost Considerations

Per-run economics are not disclosed in the source (checked 2026-09-27). Drivers an operator must budget:

  • Token volume across 12 phases × concurrent teams (interactive Latency-sensitive phases favor fast models. Evals tolerate batch)
  • Warehouse query minutes triggered by agent-generated SQL
  • Eval-framework execution (model-judged rubrics multiply inference spend)
  • Human gate time — the scarce input the whole redesign optimizes
  • Buy vs adopt: Kiro seats cost $0-$200 per user per month for 50-10,000 credits, with extra credits at $0.04 (pricing table). Spec Kit is MIT-licensed and free, but you still pay for the model usage of whichever agent runs it

Maintenance

  • Quarterly: re-run the research-provenance recipes in Commands And Recipes. Refresh the departure/CPTO status line (career facts age fastest here).
  • On any Freshworks public disclosure of Prism internals (blog, deck v2, conference talk): replace the "not disclosed" items in explanation with sourced facts and delete the corresponding unknown.