Web Search

πŸ“„ README

Why. Web search via DuckDuckGo metasearch engine β€” returns ranked results with titles, URLs, snippets. Designed to discover URLs for follow-up crawling with web_crawl. Result cache with 5-minute TTL.

How it works. Registers web_search tool. On first call, auto-creates .pi/web-search-venv/ and installs ddgs from requirements.txt. Each call writes Python script to per-call isolated temp directory (ignore/web-search/search-<random>/) β€” prevents file races under concurrent calls. Executes via pi.exec bash subprocess. Results parsed from ASCII RS (0x1E)–framed stdout. Cached in memory (5-min TTL). Declares an outputSchema and returns structuredContent: { query, returned, results } (fresh and cache-hit paths alike) plus an openWorldHint annotation (readOnlyHint is deliberately not set: the first call installs ddgs into the venv). Search-execution and parse failures return { isError: true, structuredContent: { error, query } }; empty-query and venv-setup preconditions throw for framework isError signaling. SIGTERM handled β€” Python subprocess exits cleanly with code 130 on cancellation.

Location: .pi/extensions/web-search/

Details

Architecture

β”œβ”€β”€ index.ts         # Entry: tool registration, cache, concurrency semaphore
β”œβ”€β”€ python-script.ts # Inline Python script (ddgs) as string constant
β”œβ”€β”€ executor.ts      # runSearchScript: write temp file, exec subprocess, parse delimited output
β”œβ”€β”€ venv-setup.ts    # Auto-create .pi/web-search-venv + pip install ddgs
β”œβ”€β”€ types.ts         # SearchCacheEntry, SearchResult types
└── test/            # Executor + parser tests

Execution Flow

flowchart LR
    A[tool_call: web_search] --> B[Acquire semaphore: max 5 concurrent]
    B --> C{Cache hit?}
    C -- yes, within 5min TTL --> D[Return cached results]
    C -- miss --> E[ensureWebSearchVenv]
    E --> F[write Python script to temp dir]
    F --> G[exec python3 script]
    G --> H[parseSearchResults: RS-framed payload]
    H --> I[cache in memory Map]
    I --> J[formatResults: title + URL + snippet]
    J --> K[Release semaphore, return]

Key Design Decisions

  • Per-call isolated temp directory β€” Each search writes its Python script to ignore/web-search/search-<random>/. Prevents file races under concurrent calls. Cleanup via SIGTERM handler when tool is cancelled.
  • Concurrency semaphore β€” Max 5 simultaneous searches. Prevents overwhelming DDGS API rate limits and the venv lock file.
  • Cache TTL: 5 minutes β€” In-memory Map<string, SearchCacheEntry> with timestamp-aware expiry. Stale entries pruned on each set. Cleared only on session boundaries.
  • Python subprocess with ddgs β€” Uses the minimalist ddgs Python library (DuckDuckGo search, no browser dependency). No Puppeteer/Playwright overhead. Venv auto-created on first call.
  • RS-framed output parsing β€” the Python script wraps json.dumps(...) output in ASCII RS (0x1E) bytes for reliable parsing. RS is a control character JSON escaping never emits raw, so search content cannot forge the delimiter.
  • URL encoding for markdown β€” encodeUrl() encodes parentheses () in URLs to prevent markdown link breakage.
  • SIGTERM handling β€” Python subprocess exits cleanly with code 130 on cancellation. No orphan processes.
  • Max results bounded β€” Math.min(Math.max(1, maxResults), 50).

Output Format

Search results:

1. [Title](https://example.com)
   Snippet text here...

2. [Next Title](https://example.org)
   Snippet text here...

Testing

Tests cover:

  • Search result parsing from RS-framed stdout
  • Error parsing from SEARCH_ERROR blocks
  • Empty result sets
  • Cache TTL expiry logic
  • Venv setup retry and failure modes

Copyright © 2026 SchneiderDaniel. Distributed under the MIT License.

This site uses Just the Docs, a documentation theme for Jekyll.