Methodology and maintenance
The wiki uses separate raw evidence, machine observations, and editorial synthesis. Its organization follows the persistent-wiki pattern in Karpathy's LLM Wiki note.
Evidence layers
- Each acquisition run has its own private corpus directory. It stores original public embed payloads, downloaded media, timestamps and SHA-256 hashes. Raw files and full transcripts are not published.
- FFmpeg samples a regular time grid, dense opening moments and the final frame. Every frame retains its time and asset index. Carousel photographs have image indices, not timestamps. Audio is transcribed in full where available.
- A vision model receives contact sheets and the transcript for each post and writes structured observations. These are machine notes, with model and sampling provenance, uncertainty and candidate source locators.
- Topic pages synthesize the per-post notes. They must distinguish visible detail, creator statement, guest statement and interpretation. Editorial review corrects source/attribution mistakes before strong claims are promoted.
Working with the wiki
Start at index.md, follow the relevant topic, then inspect its source pages. A citation supports only what that source shows. Do not infer off-camera habits, ownership or friendship from a single appearance. Publishing-account names and coauthor handles are textual metadata, not facial identification.
On new ingestion, preserve the previous raw corpus, collect into a new run, compare source IDs and asset hashes, update the touched topic pages, record contradictions, and append to log.md. Run the link checker after every compilation. Do not quietly replace uncertainty with a confident biography.
Interpretation limits
Visual sampling can miss fast cuts, gestures and transient objects. ASR can mishear names, slang and mixed Ukrainian/Russian speech. Automatic transcription does not diarize the host and guest reliably. It does not constitute analysis of musical structure, timbre or emotional prosody. The public discovery archive is selective and may omit most historical posts. These limits remain visible in the coverage ledger.
The initial edition was text only. The owner-requested depth revision adds the explicitly selected source-image atlas described below; it does not publish the raw media corpus or generate new video.
This edition: actual processing and review
The run used gpt-5.4 for per-post visual/transcript analysis and topic synthesis, and whisper-1 for the 47 published-run audio transcriptions. A local faster-whisper test was also performed during setup. An available Anthropic credential was rejected; no Opus analysis is claimed.
Visual sampling used a three-second grid plus 0.5, 1 and 2 seconds near the opening and a final frame. There are 1,467 sampled video frames and 163 still images. FFprobe measured 64.74 minutes across the 48 video files; Instagram metadata summed to 64.52 minutes. The small difference is a metadata/duration discrepancy, not missing analysis. One video had no audio stream.
All 66 downloaded posts received structured model analysis. These comprise 55 posts published by the target account and 11 posts with explicit target-account coauthor metadata. Their associated media comprise 47 standalone video posts and 19 carousels; one carousel contains an additional video. The public archive's dated records span May 2024–September 2026, with uneven coverage and some recently discovered posts lacking verified dates.
Editorial visual spot checks examined representative interview, car, studio, graphic-layout and lifestyle frames. They corrected layout interpretation, attribution and two mismatched reference rows. This was not an independent human review of every frame or every machine statement. The correction ledger records the consequential changes. Automated checks verify Markdown links and the reference catalogue's actual asset types and sampled timestamps; semantic correctness still requires visual review.
Acquisition attempts beyond the collector
The direct Instagram profile redirected to login. The public profile-info route returned HTTP 429; attempted feed/GraphQL routes returned HTTP 401. Public profile and individual-post embeds remained partly accessible. The collector combined those embeds with two pages of a public archive and a recent-post listing. A site-wide advertisement and an unrelated account's post were excluded rather than counted as target content.
Alternative public profile routes did not expose the full grid. A related YouTube channel catalogue was found, but full-video/subtitle retrieval did not succeed in this run. No YouTube full episodes are included in the 65-minute analysis total. No usable authorized Instagram session was supplied or found in the inspected browser setup.
Reproduce and extend
The repository contains the instadna command with collect, analyze, compile, lint, build and end-to-end run commands. The README documents their configuration. A new profile starts a separate private corpus. Public access and archive availability vary by account; the example's results are not a guarantee for another URL.
For this edition the reference check is instadna lint --wiki wiki/tymoshenko1 --root <private-corpus-directory>. Keep the private corpus and API/deployment credentials outside the site output. The site build publishes only selected wiki text, its small evidence inventory and static presentation assets; the default protected build authenticates every route. On 2026-10-01 the owner explicitly requested this edition without a login or password, so its deployed text wiki uses the public build. Raw media, full transcripts and credentials remain outside publication.
Visual and verbal depth pass · 2026-10-01
Owner feedback requested screenshots and concrete word/phrase/entity analysis, superseding the initial text-only edition. The visual atlas contains 25 newly extracted native-dimension video frames and seven re-encoded carousel photographs, all inspected for the claims in their reference cards. Web thumbnails are separate. A source hash, output hash, asset index and video time or still-image status accompany each selected reference. No generated image substitutes were used.
The language ledger searches 25 disclosed regex patterns across the original captions and 47 Whisper transcripts, preserving source/time locators without publishing full transcripts. Nine recordings received a second ASR pass with anonymous local speaker labels. Review found translation drift, fragmented labels and uncertain words; these are recorded in the ledger. No coined term, acoustic voice profile or complete all-account phrase frequency is claimed. Collection coverage remains 66 posts.