Add a Data Sharing page showing organizations searching Flock data #8

Open
mike wants to merge 24 commits from data-sharing into master
Owner

Easthampton's Flock camera is part of a nationwide network, so
external organizations can search Easthampton's data. This adds the
data pipeline and a frontend page to make those search counts visible
to residents, porting logic already used for a companion blog post.

Data pipeline

  • Add data-preprocess/summarize-network-org-searches.py, which counts
    deduplicated Network-Audit searches by requesting organization. It
    reuses lib/csv_merge.merge for ID-based dedup instead of
    reimplementing it, promoting csv_merge._column_index to public
    column_index so both modules can look up columns the same way.
  • Extract input_files_from_glob into its own module since two CLI
    scripts now need the identical glob-handling logic.
  • Wire the new script into flake.nix: a flockNetworkOrgSearchCounts
    derivation, packages.flock-network-org-search-counts,
    packages.check-flock-network-org-search-counts, and
    nix run .#update-flock-network-org-search-counts, mirroring the
    existing update-flock-requests-data pattern.
  • Verified against the source archive: 4,840 organizations and
    4,219,292 total searches, matching the public figures already
    reported for this data.

Frontend

  • Add a Data Sharing page at /data-sharing with a sortable table of
    organizations and their search counts, reusing the existing
    flock-table/flock-sort CSS and flock-main.js rendering conventions
    rather than porting the reference blog's separate chart/table
    scripts. Export compareValues from flock-model.js so the new model
    module can reuse the same sort comparison.
Easthampton's Flock camera is part of a nationwide network, so external organizations can search Easthampton's data. This adds the data pipeline and a frontend page to make those search counts visible to residents, porting logic already used for a companion blog post. Data pipeline - Add data-preprocess/summarize-network-org-searches.py, which counts deduplicated Network-Audit searches by requesting organization. It reuses lib/csv_merge.merge for ID-based dedup instead of reimplementing it, promoting csv_merge._column_index to public column_index so both modules can look up columns the same way. - Extract input_files_from_glob into its own module since two CLI scripts now need the identical glob-handling logic. - Wire the new script into flake.nix: a flockNetworkOrgSearchCounts derivation, packages.flock-network-org-search-counts, packages.check-flock-network-org-search-counts, and nix run .#update-flock-network-org-search-counts, mirroring the existing update-flock-requests-data pattern. - Verified against the source archive: 4,840 organizations and 4,219,292 total searches, matching the public figures already reported for this data. Frontend - Add a Data Sharing page at /data-sharing with a sortable table of organizations and their search counts, reusing the existing flock-table/flock-sort CSS and flock-main.js rendering conventions rather than porting the reference blog's separate chart/table scripts. Export compareValues from flock-model.js so the new model module can reuse the same sort comparison.
Add a Data Sharing page showing organizations searching Flock data
All checks were successful
ci/crow/pr/build Pipeline was successful
7a3f6c47b0
Easthampton's Flock camera is part of a nationwide network, so
external organizations can search Easthampton's data. This adds the
data pipeline and a frontend page to make those search counts visible
to residents, porting logic already used for a companion blog post.

Data pipeline
- Add data-preprocess/summarize-network-org-searches.py, which counts
  deduplicated Network-Audit searches by requesting organization. It
  reuses lib/csv_merge.merge for ID-based dedup instead of
  reimplementing it, promoting csv_merge._column_index to public
  column_index so both modules can look up columns the same way.
- Extract input_files_from_glob into its own module since two CLI
  scripts now need the identical glob-handling logic.
- Wire the new script into flake.nix: a flockNetworkOrgSearchCounts
  derivation, packages.flock-network-org-search-counts,
  packages.check-flock-network-org-search-counts, and
  nix run .#update-flock-network-org-search-counts, mirroring the
  existing update-flock-requests-data pattern.
- Verified against the source archive: 4,840 organizations and
  4,219,292 total searches, matching the public figures already
  reported for this data.

Frontend
- Add a Data Sharing page at /data-sharing with a sortable table of
  organizations and their search counts, reusing the existing
  flock-table/flock-sort CSS and flock-main.js rendering conventions
  rather than porting the reference blog's separate chart/table
  scripts. Export compareValues from flock-model.js so the new model
  module can reuse the same sort comparison.
Limit the org search table height on Data Sharing
All checks were successful
ci/crow/pr/build Pipeline was successful
a74abc20a6
The full 4,840-row table made the page unwieldy to scroll past. Cap
it to 60% of the viewport with an internal scrollbar and a sticky
header, mirroring the existing pattern used for the requests table.
Reject network searches without organization names
All checks were successful
ci/crow/pr/build Pipeline was successful
7a0ac1e970
Fail preprocessing when a deduplicated Network-Audit search has an empty or whitespace-only Org Name, instead of silently omitting it from organization counts.
Add a monthly searches chart to Data Sharing
Some checks failed
ci/crow/pr/build Pipeline failed
33f44eb388
Ports the peer repo's "Searches of Easthampton Flock data" chart,
showing monthly search volume and unique searching organizations
alongside the existing per-organization table.

Data pipeline
- Add data-preprocess/lib/timestamps.py, porting new_york_date so
  Search Time values bucket by New York calendar date rather than UTC,
  matching the peer repo's bucketing.
- Add data-preprocess/lib/network_monthly_summary.py, which reuses
  csv_merge.merge for ID-based dedup and counts requests and
  case-folded unique organizations per month.
- Add data-preprocess/summarize-network-monthly.py and wire it into
  flake.nix as flockNetworkMonthlySummary, packages.flock-network-
  monthly-summary, packages.check-flock-network-monthly-summary, and
  nix run .#update-flock-network-monthly-summary, mirroring the
  existing update-flock-network-org-search-counts pattern.

Frontend
- Vendor Chart.js and add a bar+line chart (searches per month,
  organizations searching per month) via monthly-summary.js and
  monthly-summary-model.js, following the existing org-searches.js
  model/render split.
- Make the "Since 2025, N organizations performed M searches" intro
  sentence live: org-searches.js now fills in the organization and
  search-count totals from the data it already loads, instead of
  hardcoding numbers that would go stale.
Drop dated CSV filenames for the network summary data files
Some checks failed
ci/crow/pr/build Pipeline failed
b6d185e5e2
The Month/Org Name-keyed CSVs derive their own currency from their
data, so a dated filename only needed manual upkeep. Rename them to
network-monthly-summary.csv and network-org-searches-counts.csv, and
update every reference in flake.nix, the frontend models/tests, and
the download links on Data Sharing.
Render Data Sharing charts/tables via custom elements instead of innerHTML strings
All checks were successful
ci/crow/pr/build Pipeline was successful
80011c496c
monthly-summary.js and org-searches.js built their DOM entirely from
template-literal strings assigned to innerHTML, with a hand-rolled
escapeHtml helper to make that safe. Static structure now lives in
<template> tags (app/src/_includes/custom-elements/), and
<flock-monthly-summary>/<flock-org-searches> custom elements clone
them on connectedCallback. Dynamic content (table rows, error
messages) is built with createElement/textContent instead of string
concatenation, so escapeHtml is no longer needed.

The custom elements use light DOM (no shadow root) so the existing
global stylesheet keeps styling .flock-panel/.flock-table/etc.
unchanged. Unknown custom elements default to display: inline, so
.flock-explorer picked up an explicit display: block to preserve its
centered layout.

The <template> partials are included from base.html rather than the
data-sharing.md content itself, since data-sharing.md is run through
markdown-it and <template> isn't on CommonMark's list of recognized
block-level HTML tags.
Merge CSVs with named columns via csv.DictReader
All checks were successful
ci/crow/pr/build Pipeline was successful
3d4279bc8b
Rework the CSV merge pipeline to operate on rows as dictionaries keyed by
column name instead of positional lists. merge-local-csvs.py writes the
merged table with csv.DictWriter.

Notable behavior changes:
- skipinitialspace=True replaces manual header-whitespace stripping, so
  the deployed aggregate CSV header no longer carries spaces after commas.
  The frontend already trims headers, so it is unaffected. Only spaces
  after delimiters are skipped; tab-padded or trailing-whitespace headers
  are no longer matched.
- MissingColumnError and row-width validation are gone. A missing ID or
  License Plate column propagates as a KeyError when the first data row
  is processed. Short rows fail on their None values or pass through
  with empty cells; extra columns raise ValueError in DictWriter.
- Errors identify rows by ID value rather than row number, and duplicate
  IDs are filtered in a pass after parsing.

The deliberate merge behaviors live in named helpers: _read_header,
_canonicalized_license_plate, _canonicalized_id, and _is_duplicate.

Validation: dev-scripts/build-all-flake-targets

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Stop trimming CSV headers in the frontend
All checks were successful
ci/crow/pr/build Pipeline was successful
29224a168c
The deployed aggregate CSV now has clean column names, so the PapaParse
transformHeader trim was a no-op. The header names in the CSV are the
contract; parse them as-is.

Validation: dev-scripts/build-all-flake-targets and rendered / via
nix run .#render-page to confirm the table still populates.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Resolves the semantic conflict in the Data Sharing summary modules:
network_monthly_summary and org_search_counts now call the dict-based
csv_merge.merge, which no longer takes an excluded-columns argument or
exposes column_index. Rows are dicts keyed by column name, so a missing
Org Name or Search Time column now raises KeyError instead of
csv_merge.MissingColumnError.

Validation: dev-scripts/build-all-flake-targets passes, including the
byte-identity checks for the regenerated network summary CSVs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Merge remote-tracking branch 'origin/master' into data-sharing
Some checks are pending
ci/crow/pr/build Pipeline is running
16da3fd33e
Use descriptive organization names in tests
Some checks failed
ci/crow/pr/build Pipeline failed
abcaaa2661
Match the master branch's move from single-letter placeholder names
('Officer A', 'Org A') to descriptive dummy values in test data
(#14, #15). Applies the same style to the org search counts and
network monthly summary tests added on this branch.
pip treats a lone / as a local-path requirement pointing at the
filesystem root, so pip install -r dev_requirements.txt fails.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Commit 29224a1 removed the PapaParse transformHeader trim from
flock-main.js because the generated CSVs' header names are the
contract, but the two Data Sharing loaders reintroduced it. Both new
CSVs are generated by this repository's own scripts and pinned by the
check-flock-network-* flake checks, so parse them as-is too.

Validation: node --test app/tests/*.test.js and rendered
/data-sharing.html via nix run .#render-page to confirm the chart and
table still populate.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
org-searches-model.js duplicated flock-model.js's sortRows minus the
sortValue indirection, which falls through to row[key] for the org
table's keys anyway; both produce identical output for every
key/direction combination. Re-export the existing function and make
compareValues private again.

Validation: node --test app/tests/*.test.js.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
aria-sort only allows ascending/descending/none/other, but the sort
indicators wrote the internal asc/desc tokens, so screen readers could
not announce sort state. Map the internal direction to the ARIA
vocabulary in both the explorer table and the org searches table.

Validation: rendered / and /data-sharing.html via nix run .#render-page
and confirmed aria-sort="descending"/"none" in the DOM.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Count unique organizations case-sensitively, matching how
org_search_counts.py groups the org searches table. The source data has
no case-variant organization names, so the generated CSV is unchanged.

Validation: python tests and check-flock-network-monthly-summary, which
byte-compares the regenerated CSV against the served one.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Factor shared helpers for the flake's CSV data plumbing
Some checks failed
ci/crow/pr/build Pipeline failed
1ccf7effc3
The three generated-CSV derivations, their cmp checks, and their
update apps were copy-pasted modulo name, script, and filename. Add
mkFlockCsv, mkCsvDataCheck, and mkUpdateCsvApp helpers and instantiate
all three datasets through them, so adding a dataset is one entry per
section.

Validation: dev-scripts/build-all-flake-targets.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Integrate all Flock data
All checks were successful
ci/crow/pr/build Pipeline was successful
6eb1d3d352
All checks were successful
ci/crow/pr/build Pipeline was successful
Required
Details
This pull request can be merged automatically.
You are not authorized to merge this pull request.
View command line instructions

Checkout

From your project repository, check out a new branch and test the changes.
git fetch -u origin data-sharing:data-sharing
git switch data-sharing
Sign in to join this conversation.
No reviewers
No labels
No milestone
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
oe/flock-data-explorer!8
No description provided.