SIGN IN SIGN UP

fix: [ENG-3021] render and index <img> as first-class inline content

Closes the silent-strip bug where curating a topic with `<img src alt/>`
embedded inside a `<bv-*>` element succeeded on disk but vanished on
`brv read` and was missed by `brv query`. Empirically reproduced
2026-05-29: writer kept the tag intact; `brv read` showed a gap where
the image was; query for the alt text returned zero matches.

Cause: `html-renderer.ts` and the BM25 indexer's `bodyText` extraction
both rely on `getInnerText()`, which walks text-node descendants only.
Void elements like `<img>` have attribute data (`src`, `alt`) but no
text children → contribute nothing. Worst-class UX (write succeeds,
content disappears).

Approach: treat `<img>` as first-class inline content.

- `html-renderer.ts` — new private `getInlineMarkdown(node)` that walks
  like `getInnerText` but translates `<img>` to CommonMark `![alt](src)`.
  `renderChild` uses it for inline content; a top-level `<img>` case is
  added for the rare bv-sibling shape. Defensive on malformed input:
    * missing `src` → empty string (no broken `![alt]()` syntax)
    * missing `alt` → `![](src)` (valid CommonMark click target)
    * `]` in alt → collapsed to space (only `]` closes the alt span)
    * `)` in src → CommonMark autolink form `<src>` (parens-tolerant)
- `html-reader.ts` — new exported `extractImageContent(elements)` that
  aggregates every `<img>`'s alt + src into a space-joined string.
  Surfaced via the new `HtmlTopicRead.imageContent` field. Does NOT
  mutate `getInnerText` — separate focused helper with no surprise
  blast-radius on shared infrastructure.
- `search-knowledge-service.ts` — concatenate `parsed.imageContent`
  into the indexed-content array alongside bodyText / summary / tags /
  keywords / related. URLs go in verbatim; the BM25 tokenizer's
  whitespace/punctuation split decomposes them into useful tokens
  (host, path segments, filename, extension).
- `INDEX_SCHEMA_VERSION` bumped 6 → 7 so cached indexes built pre-fix
  invalidate on next daemon start. Previously-curated `<img>` content
  becomes searchable retroactively without a manual `brv index rebuild`.
- `system-prompt.yml` — extend the inline-HTML allowlist note to
  document `<img>` is supported.

Tests (16 new + 6 integration cases):

- `html-renderer.test.ts` (7 cases): canonical `<img>` rendering inside
  `<bv-decision>`; `![](src)` for missing alt; silent drop for missing
  src; `]` escape in alt; `)` autolink fallback in src; top-level
  `<img>` sibling rendering; URL tokens present in rendered output for
  BM25 friendliness.
- `html-reader.test.ts` (6 cases): empty topic → empty imageContent;
  single image alt + src aggregated; multiple images preserve document
  order; empty attrs don't produce double spaces; `readHtmlTopicSync`
  surfaces `imageContent` on the parsed result.
- `test/integration/scenarios/img-roundtrip.test.ts` (6 cases): full
  read + index pipeline. Curate-shaped HTML on tmp disk → `readHtmlTopic`
  → renderer shows markdown image syntax; indexer (MiniSearch with the
  same option shape as production) finds the topic for queries on alt
  phrase, URL host token, URL path segment, surrounding prose. Plus a
  regression guard for topics with zero images (no double spaces in
  the indexed content).

42/42 affected-surface tests green; 54/54 search-knowledge regression
tests pass. Typecheck + lint clean; the `renderChild` complexity
warning was 31 pre-fix and is 32 now — +1 unavoidable for the new
top-level `<img>` branch.

This is the first task in the post-merge inline-html-support
milestone (`features/html-memory-conversion/milestones/02-...`).
The matching fix for `<a href>` (same shape, different bug surface)
is tracked as a follow-up task.
D
Danh Doan committed
d7b0591af1803b80a79104b299778df3f479a2cd
Parent: 99fc4e5