Skip to content

Testing ​

Patterns and tools for testing JavaScript applications across Node.js and browser environments.

Sessions started by a test are never saved ​

Code or agents that start a Claude session for a check, eval, or test must pass --no-session-persistence (print mode only; Agent SDK: persistSession: false) and run it from a temp directory, never the repo. Claude Code saves any other session under ~/.claude/projects/<the dir it ran in>/, where it shows up in that repo's session list next to real work.

  • A check that must read the repo gets it with --add-dir <repo>.
  • Set CLAUDE_CODE_DISABLE_AUTO_MEMORY=1, or each new temp directory leaves an empty folder under ~/.claude/projects/.
  • A multi-turn session keeps its conversation in one process, fed over stdin (--input-format stream-json), since nothing is saved to --resume.

eval.js and spec.js do all of this through claude-session.js.

Topics ​

  • JavaScript Testing: Testing patterns using the Node.js test runner and node:assert/strict — test exports, signatures, and observable behaviors. Use when writing or running tests in Node.js.
  • Playwright Browser Testing: End-to-end browser testing with Playwright and Chromium, including GitHub Actions CI setup. Use when tests need a real browser for navigation, clicks, or screenshots.
  • In-Browser Testing: Lightweight in-browser test runner with HTML fixtures and method stubs, plus a Playwright CI bridge. Use when tests should live in the same file as the code being tested.
  • Playwright Screenshots: Retina-quality screenshots with custom viewports and tooltip capture. Use when capturing polished screenshots for documentation or design review.
  • Playwright Mouse Cursor: Inject a visible mouse cursor overlay into Playwright pages. Use when screenshots, recordings, or demos need to show cursor position.
  • Screenshot Diff Testing: Visual regression testing that commits screenshots to git and uses git diff for pass/fail. Use when verifying visual output changes with human approval.
  • Evals: Testing agent behaviors with simulated user conversations, structured grading, and eval reports. Use when writing or running evals.
  • LLM-Judged Specs: Persistent *.spec.md assertions about the codebase, checked by an LLM judge that reads the relevant artifact and renders a verdict. Use when verifying properties that are easy to check by reading but hard to check with code (documentation freshness, convention adherence, cross-reference accuracy).