{
  "story_id": "1f3c61b23fe60bed05f6bef0eaae3c91",
  "desk": "drm3",
  "revision": 1,
  "published_at": "2026-09-03T04:00:00.000Z",
  "content_hash": "f87b7444b9f95d62873d3b49f35b8b8e4716a2f5ae5ead1ad4a17931d45554c8",
  "hash_basis": "sha256 over `headline\\ndek\\nprose`, plus `\\n` + the canonical citations JSON when any source is placed, plus `\\n#blog` for blogs",
  "basis": {
    "headline": "Vercel Builds Feedback Loop Treating Agent Instructions Like Software",
    "dek": "Vercel tested 200 agent runs to refine design.md, reducing web page generation failures by 57%.",
    "prose": "Vercel ran more than 200 agent runs to build design.md, a new public prompt file designed to help agents create web pages that look and feel like Vercel. [^1]\n\nResearchers introduced ReproRepo, a scalable framework for reproducibility evaluation that leverages human-raised GitHub issues as naturally occurring supervision on realistic reproduction blockers. [^2]\n\nCritical page-level leakage remained at 0.968 despite the improvements in localization metrics. [^3]\n\nResearchers introduced LeakageBench, a challenge set of 500 document images containing 11,954 GDPR-aligned PII annotations spanning direct identifiers, linkage keys, and contextual re-identification surfaces. [^4]\n\nThe authors present ToolGate, a system that treats every AI-generated benchmark item as a proposal subject to three validation gates. [^5]\n\nIn a test comparing Codex with GPT-5.5, Vercel's deterministic checks counted 39 instances of known failure modes with design.md loaded, compared to 91 without it. [^6]\n\nThe experiment demonstrated a 57% reduction in known failure modes when using the design.md file compared to unguided generation. [^7]\n\nVercel acknowledged that every one of the six pages tested had a failure large enough to prevent shipping, even with the guidance file. [^8]",
    "cited": "[{\"statement\":\"Vercel ran more than 200 agent runs to build design.md, a new public prompt file designed to help agents create web pages that look and feel like Vercel.\",\"source\":\"thenewstack.io\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T15:29:58.000Z\",\"publisher_count\":1,\"sources\":[\"thenewstack.io\"]},{\"statement\":\"Researchers introduced ReproRepo, a scalable framework for reproducibility evaluation that leverages human-raised GitHub issues as naturally occurring supervision on realistic reproduction blockers.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"Critical page-level leakage remained at 0.968 despite the improvements in localization metrics.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-03T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"Researchers introduced LeakageBench, a challenge set of 500 document images containing 11,954 GDPR-aligned PII annotations spanning direct identifiers, linkage keys, and contextual re-identification surfaces.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-03T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"The authors present ToolGate, a system that treats every AI-generated benchmark item as a proposal subject to three validation gates.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-03T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"In a test comparing Codex with GPT-5.5, Vercel's deterministic checks counted 39 instances of known failure modes with design.md loaded, compared to 91 without it.\",\"source\":\"thenewstack.io\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T15:29:58.000Z\",\"publisher_count\":1,\"sources\":[\"thenewstack.io\"]},{\"statement\":\"The experiment demonstrated a 57% reduction in known failure modes when using the design.md file compared to unguided generation.\",\"source\":\"thenewstack.io\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T15:29:58.000Z\",\"publisher_count\":1,\"sources\":[\"thenewstack.io\"]},{\"statement\":\"Vercel acknowledged that every one of the six pages tested had a failure large enough to prevent shipping, even with the guidance file.\",\"source\":\"thenewstack.io\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T15:29:58.000Z\",\"publisher_count\":1,\"sources\":[\"thenewstack.io\"]}]",
    "kind": "news"
  },
  "receipt_verify": "Ed25519 over the dot-joined string `slice_hash.cursor_from.cursor_to.view.view_version.row_count`; public_key and sig are base64url of the raw 32-byte key / 64-byte signature",
  "receipt": null,
  "receipt_note": "this revision predates receipt-keeping (before v0.37.0); the filed row lives in the record",
  "generation_chain": {
    "wire": {
      "stream": "fountain_news",
      "story_id": "e6d1c62ca4a6042c24311d6721b42485",
      "thread_id": "243275c650dffadc5e37eee862daa06c",
      "thread_label": "GPT-5.5",
      "novelty": "UPDATE",
      "content_hash": "5e6f67bf749c55e6c96537ca6894673a3435a8eba5c03b4ebdc0640433b5eef7",
      "last_published_at": "2026-09-03T04:00:00.000Z",
      "read_receipt": {
        "slice_hash": "a6974819026a155e1c79c99aba73ab29d5f9d0aaa9cffb05e88a08023354a690",
        "cursor_from": "eyJ0cyI6IjIwMjYtMDktMDNUMDM6MzI6MTkuMDAwMDAwWiIsImlkIjoiNzMzZDYyYTFiNGQwMmJmNjYzNTk3YjhmN2JhZDBiZTIiLCJ2IjoiMSJ9",
        "cursor_to": "eyJ0cyI6IjIwMjYtMDktMDNUMDQ6MDk6MDAuMDAwMDAwWiIsImlkIjoiNjQ5MjI3ZDNmNWQyMGQ3NmI1ZTE1NGNhODNlMTI0ZTMiLCJ2IjoiMSJ9",
        "view": "v_fountain_news",
        "view_version": "1",
        "row_count": 100,
        "window_days": 3,
        "bytes_scanned": 12568115,
        "credits": 8,
        "price_per_100_rows": 8,
        "sig": "k4fkL1dHRNIJwSawY8K6wWBrTSdTIRM-buyZr56gcruiCoH5_IdJUeMzPH7d51ZTXAE0Qe6xSFoDys5UGTzsAw",
        "public_key": "bMUigy8O0jOnBxQ4Sc-5lwhIZ8LQVAhxMbR7qESVuUE",
        "signer_path": "lakehouse/data-extract/v1",
        "alg": "Ed25519",
        "signed": true
      }
    },
    "written_at": "2026-09-03T06:46:30.909Z"
  },
  "cited_facts": [
    {
      "statement": "Vercel ran more than 200 agent runs to build design.md, a new public prompt file designed to help agents create web pages that look and feel like Vercel.",
      "source": "thenewstack.io",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T15:29:58.000Z",
      "publisher_count": 1,
      "sources": [
        "thenewstack.io"
      ]
    },
    {
      "statement": "Researchers introduced ReproRepo, a scalable framework for reproducibility evaluation that leverages human-raised GitHub issues as naturally occurring supervision on realistic reproduction blockers.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "Critical page-level leakage remained at 0.968 despite the improvements in localization metrics.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-03T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "Researchers introduced LeakageBench, a challenge set of 500 document images containing 11,954 GDPR-aligned PII annotations spanning direct identifiers, linkage keys, and contextual re-identification surfaces.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-03T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "The authors present ToolGate, a system that treats every AI-generated benchmark item as a proposal subject to three validation gates.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-03T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "In a test comparing Codex with GPT-5.5, Vercel's deterministic checks counted 39 instances of known failure modes with design.md loaded, compared to 91 without it.",
      "source": "thenewstack.io",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T15:29:58.000Z",
      "publisher_count": 1,
      "sources": [
        "thenewstack.io"
      ]
    },
    {
      "statement": "The experiment demonstrated a 57% reduction in known failure modes when using the design.md file compared to unguided generation.",
      "source": "thenewstack.io",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T15:29:58.000Z",
      "publisher_count": 1,
      "sources": [
        "thenewstack.io"
      ]
    },
    {
      "statement": "Vercel acknowledged that every one of the six pages tested had a failure large enough to prevent shipping, even with the guidance file.",
      "source": "thenewstack.io",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T15:29:58.000Z",
      "publisher_count": 1,
      "sources": [
        "thenewstack.io"
      ]
    }
  ],
  "note": "A signature proves who filed this and that it has not changed since. It never makes a claim true."
}