Unify analysis pipeline around the TSV; move tracked DBs out of cloud sync

- Tracked DBs now live at /mnt/data/projects/cupido/tracked/ (out of
  ownCloud to avoid sync conflicts and bandwidth churn). config.py
  TRACKING_OUTPUT_DIR points there; the docker-compose for ethoscope-lab
  mounts it world-readable for JupyterHub users.
- New scripts/export_video_db_index.py joins all_video_info_merged.xlsx
  with the video inventory and the on-disk DBs, producing a TSV that has
  one row per fly/ROI plus training/testing video and DB paths. Handles
  approximate xlsx times, cross-day training/testing, the 12 AM/PM
  ambiguity, and date typos.
- scripts/load_roi_data.py rewritten as a TSV-driven loader returning a
  single DataFrame with session and metadata columns. calculate_distances
  and the two flies_analysis notebooks migrated to use it; downstream
  trained/naive splits remain available via simple equality filters.
- Metadata vocabulary canonicalized: {naïve, niave, untrained, test} all
  resolve to {trained, naive}. Normalization happens at the TSV-export
  boundary (idempotent); the xlsx and the 2025-07-15 legacy CSV were
  edited in place to remove the worst variants.
- scripts/monitor_tracking.py rate calculation fixed: with N parallel
  workers, completions arrive in bursts; the old formula divided by burst
  width and reported nonsense rates. Now uses a 6 h window denominator.
- scripts/track_videos.py: BGRMovieCamera retries cv2.read on transient
  NFS hiccups and a post-tracking completeness gate (≥ 90 % of expected
  duration via MAX(t) across all 6 ROIs) deletes silent partial DBs.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
Giorgio Gilestro 2026-04-30 15:20:14 +01:00
parent e4da7691d5
commit f60a9d0530
13 changed files with 569 additions and 237 deletions

View file

@ -28,7 +28,22 @@
"execution_count": null,
"metadata": {},
"outputs": [],
"source": "# Load the pre-processed data\ntrained_data = pd.read_csv(DATA_PROCESSED / 'trained_roi_data.csv')\nuntrained_data = pd.read_csv(DATA_PROCESSED / 'untrained_roi_data.csv')\n\nprint(f\"Trained data shape: {trained_data.shape}\")\nprint(f\"Untrained data shape: {untrained_data.shape}\")\nprint(f\"Trained data columns: {list(trained_data.columns)}\")\nprint(f\"Untrained data columns: {list(untrained_data.columns)}\")"
"source": [
"# Load tracking data via the unified loader (driven by all_video_info_merged.tsv).\n",
"# Reason: replaces reads of trained_roi_data.csv / untrained_roi_data.csv with\n",
"# the live loader so the notebook always sees the current batch.\n",
"sys.path.insert(0, str(PROJECT_ROOT / 'scripts'))\n",
"from load_roi_data import load_roi_data\n",
"\n",
"data = load_roi_data()\n",
"trained_data = data[data['male'] == 'trained'].copy()\n",
"untrained_data = data[data['male'] == 'naive'].copy()\n",
"\n",
"print(f\"all data shape: {data.shape}\")\n",
"print(f\"Trained data: {trained_data.shape}\")\n",
"print(f\"Naive data: {untrained_data.shape}\")\n",
"print(f\"Columns: {list(trained_data.columns)}\")\n"
]
},
{
"cell_type": "markdown",
@ -418,4 +433,4 @@
},
"nbformat": 4,
"nbformat_minor": 4
}
}