# ig-harvester ships as an MIT-licensed Instagram OSINT collector that attaches to your own logged-in Chrome via CDP

The PowerShell-and-Node toolkit mines Relay JSON, semantic DOM and screenshots from a burner session you control, writes resumable SQLite caches, and computes engagement and posting analytics locally — with the author stating it is intended for OSINT research, journalism and security audits.

Canonical URL: https://freelancenews.online/news/ig-harvester-ships-as-an-mit-licensed-instagram-osint-collector-that-9a66c71b
Published: 2026-10-09T06:17:27.007Z
Updated: 2026-10-09T06:17:27.007Z
Source published: 2026-10-09T06:05:07.000Z
Event date: Not established
Review status: source-reviewed
Review method: Automated comparison against retrieved source text; not independent fact-checking.

## Report

A developer publishing as anurag-panda-dev has released ig-harvester, an MIT-licensed Instagram collection toolkit that drives a Chrome window the operator is already logged into rather than attempting to authenticate on its own. The project is documented in a dev.to write-up by its author and in the GitHub repository README, both of which describe the same architecture and command surface.

The design choice is the tool's central claim. According to the author, Instagram offers no public API for the data OSINT work needs — comment threads with reply hierarchies, complete follower lists, and carousel slides at original resolution — so existing tools either reverse-engineer the private API or drive a browser. ig-harvester takes the browser route but attaches to an existing session through the Chrome DevTools Protocol, launched by a bundled PowerShell script that enables remote debugging on port 9222. The operator logs into a burner account in that window and keeps it open while the scraper runs.

The author frames this as an operational and legal posture rather than a technical shortcut: the tool does not store cookies, solve challenges or rotate accounts, and the README states that nothing leaves the machine. That is an author claim about the software's behaviour, not an independently verified audit, and readers should treat it accordingly.

Extraction is deliberately redundant. The repository describes three sources mined in parallel: embedded Relay JSON in the page source, which is rich but occasionally absent; semantic DOM traversal, which is always present but shallower; and screenshots as ground truth when the first two disagree. The author cites a concrete failure case — Instagram recently stopped embedding sidecar JSON for carousels — and says the tool fell back to walking the post with a Next control and reading each slide's image source at full resolution.

Resumability is handled through a SQLite cache keyed by shortcode. A run writes JSON, per-post and per-comment CSVs, follower and following lists, an optional SQLite database, and a hidden resume cache; killing the process and restarting continues from the last collected item unless the operator passes a flag to disable resume. The author notes that a 500-post harvest takes a while and browsers crash, which is the stated rationale.

Pacing is configurable rather than fixed. The documented controls include jittered delays between actions with millisecond bounds, a token bucket for burst control, exponential backoff on transient failures, and optional HTTP or SOCKS5 proxy support. These are described as mechanisms in the source, not benchmarked against any detection system.

Output is not limited to raw records. Each run computes engagement rate, posting cadence, best posting hour and day, top hashtags and mentions, follower-to-following ratio, and bio signals such as email, phone or URL detection, written into the JSON payload. A separate ig-analyzer process serves a local dashboard on port 8080 charting likes per post, posts by hour and day, and content mix.

A third component, ig-images, does one job: saving every post image, including each carousel slide, as a PNG named after the post timestamp. The README states it shares the browser attach, authentication, grid collection and configuration flags with the main scraper but collects nothing else — no JSON, CSV, comments or follower lists.

Archiving is separated from scraping. A PowerShell import script copies each run into a dated snapshot under a structured directory, regenerates reports and a master index, and never overwrites older runs, so follower counts remain comparable across time. Media is stored once in a latest folder rather than duplicated per snapshot, and the README notes that an import without images preserves the existing media archive instead of deleting it.

The entry point is a single PowerShell script that runs either an interactive menu or a named action, with exit codes documented as 0 for success, 1 for a failed step or preflight, and 2 for a usage error or cancelled input. A non-interactive flag answers prompts with defaults, and a dry-run flag prints the exact commands without executing them. Preflight checks Node 20 or newer, npm dependencies and the CDP endpoint, offering to launch the debug Chrome if the endpoint is down.

A limitation that matters for anyone planning a collection is made explicit in the README: because post links are read from the profile grid, a run returns only what Instagram serves to the attached session. Both numbers appear in a stopped-early message logged by the tool when the profile header reports more posts than the grid returned, and rather than a silent short list, private accounts produce an explicit error. The same grid as any other follower is seen by a burner account that already follows a private profile.

For freelancers and developers, the practical appeal is the shape of the dependency: a local Node and Chrome setup, no API keys, no hosted service, and a documented Docker path. That lowers the barrier to running a one-off collection on a workstation, and the resumable cache and dated archive suit repeated measurement of the same public profile over time. The tradeoff is that everything depends on a live logged-in browser session, so throughput, reliability and results are bounded by what that session is shown — not by the tool's own logic.

The author's stated use cases are OSINT research, journalism, security audits and education, with an explicit note that operators are responsible for complying with applicable law, Instagram's terms and privacy regulations such as GDPR and CCPA, and for confirming consent before collecting data about someone. The README repeats that framing and disclaims liability for misuse.

Several things remain unverified in the supplied material. There is no independent test of the extraction fallbacks, no measurement of collection speed or detection rates, no pricing question because the project is free and MIT-licensed, and no stated release version or changelog in the evidence. The repository lists requirements of Node 20 or newer and Chrome, and suggests running the install and test commands to verify a setup, but the evidence does not report results from that test suite.

The most useful way to read this release is as a documented, inspectable implementation of a known approach — browser automation over a session you own — with the interesting engineering in the fallback chain and the resume cache rather than in any novel access method. Whether the triple-source extraction holds up as Instagram changes its markup is exactly the kind of claim that only repeated real-world runs can settle, and the author invites that feedback by asking how others handle extraction fallbacks.

## Key points

- ig-harvester is MIT-licensed and attaches to an existing logged-in Chrome session over the Chrome DevTools Protocol on port 9222, rather than storing cookies or authenticating itself.
- It cross-checks three extraction sources — embedded Relay JSON, semantic DOM and screenshots — and the author says it survived Instagram dropping sidecar JSON for carousels by walking slides and reading image sources.
- Runs write JSON, per-post and per-comment CSVs, follower and following lists, an optional SQLite database and a shortcode-keyed resume cache, so interrupted harvests continue instead of restarting.
- Post links come from the profile grid, so results are limited to what the attached session is served; the tool logs a stopped-early count and an explicit private-account error instead of returning a silent short list.
- The author states the tool is intended for OSINT research, journalism, security audits and education, and places responsibility for legal compliance and consent on the operator.

## Practical implications — editorial interpretation

For developers and freelancers who need structured public-profile data without a hosted scraping service, this is a self-contained local option: Node 20 plus Chrome, a burner login, and a documented CLI and Docker path. The dated archive and resume cache make repeated measurement of the same profile feasible, but the ceiling on any collection is what the attached session is shown, so plan for partial grids and log the stopped-early counts rather than assuming a complete follower or post list.

## Limitations and unknowns

All technical claims come from the author's dev.to post and the project README; no independent testing, benchmark or third-party review is present in the evidence. No version number, release date or changelog is supplied, and the repository's own test command is mentioned without reported results. The evidence does not quantify collection speed, detection risk or accuracy of the extraction fallbacks, and the privacy and compliance posture is the author's stated intent rather than a verified property of the software.

## Sources

- [1] dev.to: ig-harvester: an open-source Instagram OSINT toolkit built on Playwright + CDP
  https://dev.to/anurag_panda/ig-harvester-an-open-source-instagram-osint-toolkit-built-on-playwright-cdp-33l7
  Retrieved: 2026-10-09T06:17:01.659Z
- [2] github.com: IG-HARVESTER is an open-source (MIT) Instagram OSINT tool for collecting posts, comments, followers and media via Playwright CDP.
  https://github.com/anurag-panda-dev/ig-harvester
  Retrieved: 2026-10-09T06:17:10.928Z

## Claim references

- The toolkit attaches to an already logged-in Chrome window via the Chrome DevTools Protocol instead of managing its own authentication. [source 1]
- The author says Instagram lacks a public API for the data OSINT work needs, including comment threads, full follower lists and original-resolution carousel slides. [source 1]
- Extraction combines embedded Relay JSON, semantic DOM traversal and screenshots, with screenshots used as ground truth when the other two disagree. [source 1]
- The author reports that when Instagram stopped embedding sidecar JSON for carousels, the tool fell back to walking the post and reading each slide's image source at full resolution. [source 1]
- Post links are read from the profile grid, so a run returns only what Instagram serves to the attached session, and the tool logs a stopped-early count when the header exceeds the grid. [source 2]
- The tool is MIT-licensed and the author states it is intended for legitimate OSINT research, journalism, security audits and education, with compliance responsibility on the user. [source 2]
