{
  "story_id": "d9b197ed7a35f29c3ccbd70ea4494e43",
  "desk": "drm3",
  "revision": 1,
  "published_at": "2026-09-04T04:00:00.000Z",
  "content_hash": "b00fb23d0f0e8ef7b2a7e924fa3c46c22718f214e7962faf8ea37ded8d9abb14",
  "hash_basis": "sha256 over `headline\\ndek\\nprose`, plus `\\n` + the canonical citations JSON when any source is placed, plus `\\n#blog` for blogs",
  "basis": {
    "headline": "TAME Framework Improves Text-Video Retrieval Using Temporal Modeling",
    "dek": "Researchers propose TAME, a CLIP-based framework with temporal modeling to enhance text-video retrieval performance.",
    "prose": "Researchers propose Temporal-Aware Mixture-of-Experts for Text-Video Retrieval (TAME), a CLIP-based framework that jointly models frame-level structure and temporal relations. [^1]\n\nText-Video Retrieval (TVR) is fundamentally limited by the lack of temporal modeling when extending image-text models like CLIP to videos. [^2]\n\nThe paper's authors report that BioCLIP2 achieved 72.36% accuracy on the BFF-15 dataset using English common names and 68.91% on the SylFishBD dataset using scientific names. [^3]\n\nThe TAME framework achieves consistent gains on DiDeMo, MSVD, LSMDC, and ActivityNet benchmarks compared to CLIP-based baselines. [^4]\n\nVideos exhibit frame-wise heterogeneity in appearance and motion, and compressing all frames into a single representation often obscures temporal structure and semantic transitions. [^5]\n\nThe paper concludes that zero-shot biological vision-language model scores jointly reflect biological specialization, multilingual alignment, nomenclature, prompt formulation, and context. [^6]\n\nThe paper's authors report that BioCLIP2 with Bengali prompts performed near chance with balanced accuracy between 14.22% and 14.29%, while Jina CLIP partially recovered Bengali discrimination to 21.89% and 16.36% on the two sources, and bare Bengali names returned to 14.29% on both. [^7]\n\nThe paper's authors report that generic CLIP scored 25.15% on BFF-15 and 14.40% on SylFishBD, substantially lower than BioCLIP2. [^8]",
    "cited": "[{\"statement\":\"Researchers propose Temporal-Aware Mixture-of-Experts for Text-Video Retrieval (TAME), a CLIP-based framework that jointly models frame-level structure and temporal relations.\",\"source\":\"takara.ai\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T07:12:50.000Z\",\"publisher_count\":1,\"sources\":[\"takara.ai\"]},{\"statement\":\"Text-Video Retrieval (TVR) is fundamentally limited by the lack of temporal modeling when extending image-text models like CLIP to videos.\",\"source\":\"takara.ai\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T07:12:50.000Z\",\"publisher_count\":1,\"sources\":[\"takara.ai\"]},{\"statement\":\"The paper's authors report that BioCLIP2 achieved 72.36% accuracy on the BFF-15 dataset using English common names and 68.91% on the SylFishBD dataset using scientific names.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-04T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"The TAME framework achieves consistent gains on DiDeMo, MSVD, LSMDC, and ActivityNet benchmarks compared to CLIP-based baselines.\",\"source\":\"takara.ai\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T07:12:50.000Z\",\"publisher_count\":1,\"sources\":[\"takara.ai\"]},{\"statement\":\"Videos exhibit frame-wise heterogeneity in appearance and motion, and compressing all frames into a single representation often obscures temporal structure and semantic transitions.\",\"source\":\"takara.ai\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T07:12:50.000Z\",\"publisher_count\":1,\"sources\":[\"takara.ai\"]},{\"statement\":\"The paper concludes that zero-shot biological vision-language model scores jointly reflect biological specialization, multilingual alignment, nomenclature, prompt formulation, and context.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-04T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"The paper's authors report that BioCLIP2 with Bengali prompts performed near chance with balanced accuracy between 14.22% and 14.29%, while Jina CLIP partially recovered Bengali discrimination to 21.89% and 16.36% on the two sources, and bare Bengali names returned to 14.29% on both.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-04T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"The paper's authors report that generic CLIP scored 25.15% on BFF-15 and 14.40% on SylFishBD, substantially lower than BioCLIP2.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-04T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]}]",
    "kind": "news"
  },
  "receipt_verify": "Ed25519 over the dot-joined string `slice_hash.cursor_from.cursor_to.view.view_version.row_count`; public_key and sig are base64url of the raw 32-byte key / 64-byte signature",
  "receipt": null,
  "receipt_note": "this revision predates receipt-keeping (before v0.37.0); the filed row lives in the record",
  "generation_chain": {
    "wire": {
      "stream": "fountain_news",
      "story_id": "4f5b3cc1932b56014cb8ba9f2fd81343",
      "thread_id": "a45e5a4ecb73635333bd308af699c7e3",
      "thread_label": "CLIP",
      "novelty": "UPDATE",
      "content_hash": "58ff2ec480994be8a33bf5b17f9fe28544b9747e9f7eb2f3b7243a5451eae2f1",
      "last_published_at": "2026-09-04T04:00:00.000Z",
      "read_receipt": {
        "slice_hash": "226f85eeb969ef4e22c236c02a6ac83a01bfea21805771a37ad265468b4f0d1b",
        "cursor_from": "eyJ0cyI6IjIwMjYtMDktMDRUMDQ6MDA6MDAuMDAwMDAwWiIsImlkIjoiMjdkNmMxNTc0ZTI3OThkMTZkYWY2ZDc4YWYzODMzYzYiLCJ2IjoiMSJ9",
        "cursor_to": "eyJ0cyI6IjIwMjYtMDktMDRUMDQ6NDE6MTEuMDAwMDAwWiIsImlkIjoiYzI1ZDc3ZDNmMWFkMDY3MDVkMDMwZjM5NzMxYWM2YjEiLCJ2IjoiMSJ9",
        "view": "v_fountain_news",
        "view_version": "1",
        "row_count": 100,
        "window_days": 3,
        "bytes_scanned": 12166856,
        "credits": 8,
        "price_per_100_rows": 8,
        "sig": "aJ2O3M7WPo_Nx1rM5M-BmixwZpOjkJ_YQzXWupf3dFzyLFVqGAMLRh5EhYVvs8oAV9PfMSVGKw-zapLEln_zAg",
        "public_key": "bMUigy8O0jOnBxQ4Sc-5lwhIZ8LQVAhxMbR7qESVuUE",
        "signer_path": "lakehouse/data-extract/v1",
        "alg": "Ed25519",
        "signed": true
      }
    },
    "written_at": "2026-09-04T07:01:36.547Z"
  },
  "cited_facts": [
    {
      "statement": "Researchers propose Temporal-Aware Mixture-of-Experts for Text-Video Retrieval (TAME), a CLIP-based framework that jointly models frame-level structure and temporal relations.",
      "source": "takara.ai",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T07:12:50.000Z",
      "publisher_count": 1,
      "sources": [
        "takara.ai"
      ]
    },
    {
      "statement": "Text-Video Retrieval (TVR) is fundamentally limited by the lack of temporal modeling when extending image-text models like CLIP to videos.",
      "source": "takara.ai",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T07:12:50.000Z",
      "publisher_count": 1,
      "sources": [
        "takara.ai"
      ]
    },
    {
      "statement": "The paper's authors report that BioCLIP2 achieved 72.36% accuracy on the BFF-15 dataset using English common names and 68.91% on the SylFishBD dataset using scientific names.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-04T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "The TAME framework achieves consistent gains on DiDeMo, MSVD, LSMDC, and ActivityNet benchmarks compared to CLIP-based baselines.",
      "source": "takara.ai",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T07:12:50.000Z",
      "publisher_count": 1,
      "sources": [
        "takara.ai"
      ]
    },
    {
      "statement": "Videos exhibit frame-wise heterogeneity in appearance and motion, and compressing all frames into a single representation often obscures temporal structure and semantic transitions.",
      "source": "takara.ai",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T07:12:50.000Z",
      "publisher_count": 1,
      "sources": [
        "takara.ai"
      ]
    },
    {
      "statement": "The paper concludes that zero-shot biological vision-language model scores jointly reflect biological specialization, multilingual alignment, nomenclature, prompt formulation, and context.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-04T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "The paper's authors report that BioCLIP2 with Bengali prompts performed near chance with balanced accuracy between 14.22% and 14.29%, while Jina CLIP partially recovered Bengali discrimination to 21.89% and 16.36% on the two sources, and bare Bengali names returned to 14.29% on both.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-04T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "The paper's authors report that generic CLIP scored 25.15% on BFF-15 and 14.40% on SylFishBD, substantially lower than BioCLIP2.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-04T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    }
  ],
  "note": "A signature proves who filed this and that it has not changed since. It never makes a claim true."
}