Second Brain Link Documentation
Turn your data exports — 24 sources, from LinkedIn, Facebook, Instagram, Google, X and GitHub to your company's Slack, Notion, Jira and Salesforce — into a private, local, AI-queryable second brain. This page covers everything: install, build, reference, and how it all works.
Introduction
Your life is scattered across platforms — connections on LinkedIn, friends on Facebook, follows on Instagram, contacts and calendar in Google, code on GitHub, taste on Spotify. Each gives you a data export, and each sits dead in a zip. Second Brain Link pulls them into one structured knowledge vault your AI can think with — and the same person across two networks becomes a single, richer note.
- You download your own data archive(s) — from one network or several. They're yours; each platform provides them on request.
- Second Brain Link runs 100% locally, detects which source(s) you've given it, and transforms them into one clean Obsidian vault: people, companies, your voice, your career intent, your interests, and how the algorithms categorize you.
- You point Claude (or Codex) at the vault. Now you have a digital twin that knows your whole history across platforms and can reason, plan, and draft from it.
Nothing is uploaded. No account, no server, no telemetry. The output is plain Markdown files you fully own — delete it, fork it, grep it.
Why it's different
- Local-first and private by construction. The transform makes zero network calls; your data never leaves your machine.
- Privacy is in the code, not a policy. Contacts' emails and phone numbers are stripped at parse time; message bodies are never imported — only a "how often / how recently" signal.
- You own the output. Plain Markdown, no lock-in.
- MIT licensed, free, no catch.
Using the app — 3 steps
However you open Second Brain Link — the Studio web app or the desktop app — the flow is the same. A short wizard walks you through it on first launch; here it is in full.
Install the skill & run your AI locally
Add the Second Brain skill to Claude Code or OpenAI Codex and run it on your own machine — it does the building and reasoning. Cloud run — paid, coming soon
Grab the skill straight from the repo: Claude skill · OpenAI skill.
Add your sources & (re)seed the vault
Download your data exports — any of the 24 supported sources (LinkedIn, Facebook, Instagram, Google, X, GitHub, Spotify… or company-side Slack, Notion, Salesforce…) — then upload them in Studio and run the build. Studio seeds — or re-seeds — your vault from the fresh data, fully on your machine.
Connect your assistant
Use Claude Code / Codex locally, or paste your Anthropic / OpenAI / … API key in Studio's Assistant settings. Either way, calls go straight to the provider — your keys and notes never pass through us.
The local, open-source flow is free forever. Hosted cloud seeding & vault storage are an optional paid add-on — see pricing.
Second Brain Studio — web & desktop
Studio is the friendly front door to your brain: a live workspace where you browse the knowledge graph, read & edit notes, and chat with an AI that reasons over everything. It runs locally — your API keys and data stay on your machine.
Open the web app
- Go to the Studio web app — nothing to install. It opens your vault right in the browser.
- First run shows the 3-step wizard above. Drop in your export zips to seed the vault, then open the graph, notes, or chat.
Install the desktop app
- Download for your platform: macOS · Windows.
- The desktop app is the same Studio in a native window — better for working offline and keeping a vault open alongside Obsidian.
Where to go for what
Vault & Graph
Explore people, orgs, places and posts. Filter, search, and click any node to open its note.
Assistant (chat)
Ask in plain language; it cites the notes it used. Add your API key in Assistant settings, or route through local Claude Code / Codex.
Sources
Upload new export zips any time and re-run the build to refresh your vault with the latest data.
Second Brain Mobile — iOS & Android
The native mobile app is the whole pipeline in your pocket: it unpacks your exports and builds the brain entirely on-device — no account, no backend, nothing uploaded. Explore the knowledge graph, browse your sources, and talk to your twin wherever you are.
Get the app
- Download on the App Store (iPhone) · Get it on Google Play (Android) — on the App Store and Google Play.
- Onboarding walks you through it: pick your sources → import the export zips on the phone → the brain builds locally in about a minute.
Where to go for what
Twin (chat)
Talk to your brain in plain language — warm intros, meeting prep, reconnects. On iPhone the AI runs on-device via Apple Intelligence; on Android (or older iPhones) bring your own key, stored in the phone's keychain.
Brain (graph)
A native, GPU-rendered 3D graph of everyone you know and everything they connect to — search it, filter it, tap into any node.
Sources
Import LinkedIn, Facebook, Instagram and Google exports on the phone; add a newer archive any time to refresh. Everything stays on the device.
Privacy: the mobile app has no account and no cloud — the brain is built and stored only on your phone. If you use a cloud AI key, only your question and the few notes it cites go to that provider, directly.
Quick start
From zips to a working brain in about a minute of compute. You'll need Claude Code or OpenAI Codex, Python 3.8+, and (recommended) Obsidian. No other dependencies.
Download your data ~5 min per platform
Each platform hands you an export on request — pick JSON where asked:
- Settings & Privacy → Data Privacy
- Get a copy of your data
- Choose the larger archive
Can take up to 24h · link valid ~72h
- Settings → Your information
- Download your information
- Format: JSON
- Accounts Center → Your information and permissions
- Download your information
- Format: JSON
- Open takeout.google.com
- Deselect all
- Select Contacts, Calendar, YouTube, Maps, Profile
These four are the richest starting points — but the engine reads 24 sources: X, WhatsApp, GitHub, YouTube, Strava, Reddit, Spotify and TikTok on the personal side, plus 12 company sources from LinkedIn Page and Slack to Notion, Jira, Salesforce and Zendesk. Drop in whatever you have; it detects and merges them all.
Organize by entity 1 folder per person/company
Rename your-name to your actual name — the folder name becomes your brain — then unzip each export into its source subfolder. One source is fine; several get merged.
Personal + company exports in one run build sibling vaults — plus a _correlations/ brain linking people across them. Everything under data/ is git-ignored.
Install the Agent Skill one command
Run this
Restart your agent — the second-brain-link skill loads at startup. Verify with ls ~/.claude/skills/second-brain-link/.
Claude.ai / Claude Desktop: upload dist/claude/second-brain-link.skill under Settings → Capabilities → Skills (enable Code execution).
Build your brain ~60 s · zero AI cost
Say this to your agent
"Build my second brain from the exports in data/ — profile them first, show me the planned structure, then build with --full."
Building a Company Brain instead? Same thing:
"Build our Company Brain from the exports in data/company/acme/ — and emit the GBrain format too."
Under the hood the skill runs the deterministic engine for you — zero AI tokens for the build itself:
Open it and start asking Obsidian · Claude Code · Codex
- Open the vault folder in Obsidian — start at
Home.md;_STRUCTURE.mdmaps every folder; Graph view colors every note by source. - Work it inside Obsidian with the Claudian plugin (Claude Code in a side panel) or Obsidian Copilot (vault-wide Q&A + generated
/commands). - Or point Claude Code / Codex at the folder from a terminal and give it a goal — e.g. "who are my warmest paths to a Series A investor?"
Supported sources
Give it one source or a combined archive with several — it detects and merges them all. Unknown files are summarized, never dropped. For the full connector directory — what each source contributes, where it maps in the brain, and the roadmap — see the Sources page.
Personal sources — a digital twin
Profile, positions, skills, connections, companies, recommendations, posts/comments, job applications, algorithmic inferences, ad profile, search history.
- linkedin.com → Me → Settings & Privacy → Data privacy
- Get a copy of your data → choose the larger archive (everything)
- Partial archive in ~10 min; the full one within 24 h — use the second email's link
Import folder: linkedin · download link valid ~72 h
Friends → people, posts/comments → voice, liked pages → interests, ad-interests/predictions → the algorithmic mirror, check-ins → places, message signal (no bodies).
- accountscenter.facebook.com → Your information and permissions
- Download your information → Facebook profile
- Format JSON · media quality low · range All time — ready within hours
Import folder: facebook
Profile, followers/following, posts & captions, topics/interests, message signal.
- Same Accounts Center flow — choose the Instagram profile
- Format JSON
Import folder: instagram
Contacts → people, Calendar → events, YouTube subscriptions → interests, Maps → places, Profile → identity.
- takeout.google.com → Deselect all
- Select Contacts, Calendar, Maps (your places), Saved, Location History, My Activity, Chrome, YouTube, Photos
- Export once — large exports arrive as several
Takeout Nzips; drop them all in together
Import folder: google
Tweets & replies → voice, follows → people, likes → interests, lists, note-tweets, DM signal, ad data → the mirror.
- x.com → Settings → Your account → Download an archive of your data
- Re-verify, then wait 24–48 h for the email → download the ZIP
Import folder: x — unzip so the data/*.js files are inside
Exported chats → per-person message signal (frequency + recency grades relationship strength). Message text is never read.
- Per chat: open chat → ⋮ → More → Export chat → Without media
- Save the
.txt— repeat for the chats that matter
Import folder: whatsapp
Repos → projects, stars → interests, followers/following & orgs → people and organizations, issue/PR activity.
- github.com → Settings → Account → Export account data → Start export
- Email link arrives in hours — extract the tar.gz first
Import folder: github
Subscriptions & playlists → interests, watch/search history → the algorithmic mirror, your comments → voice.
- takeout.google.com → Deselect all → YouTube and YouTube Music only
- In options keep history / subscriptions / playlists / comments — untick videos
Import folder: youtube · inside a full Takeout the google source handles it
Activities → places & routines, clubs → organizations, followers → people — where you actually train.
- strava.com → Settings → My Account → Download or Delete Your Account
- Download Request → email ZIP
Import folder: strava
Posts & comments → voice, subreddits & multireddits → interests, friends → people, PM signal.
- reddit.com/settings/data-request → full GDPR export
- Email link — typically <48 h, up to 30 days
Import folder: reddit
Listening history & library → taste, playlists, artist follows, marketing inferences → the mirror. Tick "Extended streaming history".
- spotify.com/account → Security & privacy → Download your data
- Tick Extended streaming history (the default alone is thin)
- Email links (account ~5 days, extended up to 30) — drop both zips' contents together
Import folder: spotify
Follows → people, likes & hashtags → interests, watch/search history → the mirror, your comments → voice.
- App → Profile → ☰ → Settings → Account → Download your data
- Format: JSON (machine-readable) — ready in days; download within 4 days
Import folder: tiktok
Company sources — a Company Brain
Org profile, employees, followers, company posts.
- Page admin → Analytics / Settings
- Export the employee list, followers, posts and analytics reports (per-report CSVs — no single ZIP)
Import folder: linkedin_company
Directory users → employees, shared calendars → events, shared drives → projects.
- admin.google.com → Directory → Users → Download users (CSV)
- Calendars / drives via org Takeout or Google Vault
Import folder: google_workspace
Members → people, channels → projects, messages → signal only (no bodies).
- Workspace admin →
<workspace>.slack.com/services/export→ Export (public channels; all paid tiers) - Private channels / DMs need an Enterprise-Grid compliance export
Import folder: slack
Workspace export → knowledge pages & databases, page authors → people, teamspaces → structure.
- Workspace owner → Settings → Export all workspace content
- Markdown & CSV, include subpages → email link (large spaces take hours)
Import folder: notion
Space exports → the knowledge layer, page authors & contributors → people.
- Space admin → Space settings → Export space
- XML (full fidelity; HTML also parses) → ZIP
Import folder: confluence
Issues → projects, components & ownership, assignees/reporters → people — answers "who owns this?".
- Issue navigator → filter the project(s) → Export → CSV (all fields)
- Repeat per 1000-row page if large
Import folder: jira
Accounts → customers, contacts & leads → people, opportunities → the pipeline (one note per deal), campaigns, cases → support.
- Setup → Data Export Service → Export Now (or schedule)
- Email link — expires in 48 h → CSV ZIP(s)
Import folder: salesforce
Contacts, companies, deals & tickets → people, customers, pipeline and support — with owners preserved.
- Per object: Contacts / Companies / Deals / Tickets list → Export → CSV, all properties
- Keep the
hubspot-…filenames
Import folder: hubspot
Organizations & requesters → customers and people, ticket signal → support themes.
- Admin Center → Account → Requests to export data (Zendesk support enables it once)
- JSON/CSV → email
Import folder: zendesk
Any .mbox archive → people & message signal from headers. Bodies are never read.
- Gmail: Takeout → Mail → MBOX (or Google Vault org-wide)
- Outlook PST: convert first —
readpst -o out/ archive.pst→ drop the.mbox
Import folder: email — a personal Takeout containing Mail is deliberately NOT auto-imported; place the mbox here on purpose
Purview eDiscovery results → directory people + mail signal, headers only.
- Purview compliance portal → eDiscovery → Content search
- Export results (needs the eDiscovery Manager role) —
.eml+ results CSV
Import folder: microsoft365
Purview message report → per-person collaboration signal. Bodies are never read.
- Purview eDiscovery message report (CSV/JSON) for the teams in scope
Import folder: teams
That's 24 sources — 12 personal, 12 company — every one detected automatically. Anything else with a data export is usually just a JSON mapping away, and files no source claims are rescued by the universal harvester, summarized in _COVERAGE.md — never dropped.
How it works
Every network changes its export format over time and across regions. A hardcoded converter breaks the moment that drifts — so Second Brain Link detects, adapts, builds, then verifies it left nothing behind:
Which source(s)? Reads every file & column → schema map.
Optimizes the mapping to your actual export schema.
Normalizes all sources into one vault, in seconds.
Coverage report; loops if gaps remain.
- Detect + profile. Identifies the source(s) in your archive, reads every file and column, and writes a
schema_map.md— each file classified, each column typed, with privacy-safe samples. It also draws a mindmap and designs the brain structure. - Adapt. Each source is a declarative JSON mapping (
engine/mappings/sources/<name>.json). Files no mapping claims are rescued by a universal harvester and flagged in_COVERAGE.mdunder "Needs a mapping". Only genuine unknowns need your AI — it reads the schema map (never your raw data) and proposes a mapping. - Build. A deterministic, dependency-free Python engine writes the entire unified vault in seconds with zero AI cost, merging people who appear in more than one source.
- Verify (+ correlate).
_COVERAGE.mdreports how much of your export was used;_SUMMARY.mdgives seed counts. Anything unrecognized lands in99-uncategorized/— nothing is silently dropped. Multiple entities get a_correlations/brain linking them. Once built, asking the brain runs graph-grounded retrieval — see How the AI answers.
Self-adapting: the skill fixes itself
Platform exports drift — new files, renamed columns, regional variants, format changes between the export you took last year and the one you took today. The skill is built to absorb that drift on your machine, using the AI only on the residual:
- Schema map first. The profiler reads every file and column in your archive and writes
schema_map.md(PII-safe — column names only). Known shapes route to sources; unknowns are flagged, never guessed at. - Declarative mappings, not hardcoded parsers. Most sources are JSON mappings under
mappings/sources/. When your export has an extra field or a renamed file, the AI extends the mapping — or drops amapping_overrides.jsonbeside your data — instead of waiting for a code release. A mapping override wins over the built-in adapter. - The harvester rescues the rest. Files no source claims go through a universal shape recognizer (people / places / posts / interests / message signal), so data lands in the brain even before a proper mapping exists. They're listed in
_COVERAGE.mdunder "Needs a mapping" — the self-improvement backlog. - Self-heal on failure. If a build errors, the engine writes a structured
_ERROR.mdnaming the script, the file and the line. The AI applies the smallest fix, mirrors it, and re-runs — no human debugging session required. - The ratchet. Every adjustment is durable: the extended mapping keeps working on your next export, and
_DATA_POINTS.md's field-by-source matrix shows exactly which new cells lit up.
This is why the honest pitch is deterministic core, AI on the residual: the build itself costs zero tokens, and intelligence is spent only where your export genuinely differs from the known world.
Identity resolution
If "Sam Patel" is a LinkedIn connection, a Facebook friend, and a Google contact, you don't get three notes — you get one, tagged sources: [linkedin, facebook, google], with company and role filled in from whichever source knew it.
People who appear across multiple networks are surfaced as your strongest multi-context relationships. Matching is conservative by design — a wrong merge is worse than a miss (sharper fuzzy matching lands in v1.2).
The vault it builds
A layered, Obsidian-native vault. _STRUCTURE.md is the map — it documents what every folder and file is, so you (or an agent) always know where things live.
Obsidian-native by construction
- The graph is real. Filenames equal titles, so
[[Acme Cloud]]resolves to the actual note — open Graph view and you literally see your network. - Properties work. Valid YAML frontmatter with native Properties:
type,strength, ISO dates; wikilinks in properties are quoted so backlinks resolve. - Tags & filters. Every note carries
source/<name>+ type tags so the graph and search filter cleanly by where each fact came from. - One brain per entity. Multiple identities/companies each get their own vault, plus a
_correlations/brain withworks_atedges between them.
Owner mode (--full): the default build is privacy-safe. For your own brain on your own machine, --full keeps everything — emails, phones, every extra column — still 100% local.
How the AI answers — spreading activation
When you ask your brain a question, the query flows through the graph, not just over its text. Retrieval is deterministic, client-side, and makes zero network calls — the brain fires before the model speaks.
Lexical match: note titles & aliases (weight 1.0), body keywords (0.5).
Personalized PageRank over your typed, weighted edges — works_at, correlated, attended… Damping 0.85, 20 iterations.
Top-activated notes + the typed links among them become the context pack.
The model answers from the pack and cites the notes it used.
- Seed. Your words, matched directly against note titles, aliases and body keywords. The notes that match become the input activations — nothing hidden, nothing inferred.
- Spread. Activation propagates along your real links — the typed, weighted edges the engine built from your exports. That's why an old note two hops away from your question still surfaces: the graph carries it there. This is spreading activation, the textbook cognitive-science model of memory recall, computed as Personalized PageRank.
- Pack. The top-activated notes, plus the typed relationships among them, become a plain-text context pack — something you could open and read yourself.
- Answer. The model must answer from the pack, and the answer cites the notes it used. No pack, no claim.
Why you can trust the animation
The Neural view — the layered view, your vault's layers as columns — animates the wave, and every number in it is real: seed pulses, arrival order and brightness all come from the retrieval run. An edge only lights if it genuinely connects two activated notes. No fake cascades, no fake ML.
The anatomy it fires through
Studio renders every vault as tissue, at every scale — the same anatomy the wave travels:
- Cells — real sub-clusters computed from the link graph: each cell is a genuine hub note plus the notes that actually link to it. Deterministic — same vault, same anatomy.
- Membranes & tissue — soft outlines around cells and smooth blobs around regions delimit the structure without hiding anything.
- The halo — notes with no links yet hold a seat on the rim, visible until the graph grows around them.
- Cells that open — hover names a cell, click opens it in place, “only this” isolates it. Counts are always honest: showing N of M.
At 25,000 notes
Big vaults stream metadata first: the full anatomy renders and searches instantly, and note bodies load on demand the moment you open one. 100% of notes stay visible and discoverable — nothing is sampled away.
Goals & dashboards
A brain is only useful once pointed at a goal. After building, just tell your agent what you want the brain for:
"Point my brain at my goals: I'm fundraising for <thesis> and job-hunting — build the goal workspaces and color the graph."
The skill runs the analyzer for you — deterministic and re-runnable in seconds (it reads frontmatter, no rebuild):
It writes:
95-goals/<goal>.md— ranked tables answering the goal, each ending in a ready-to-run AI prompt. Built-ins: fundraising, bd, jobsearch, datamining, personalization.Dashboard.md— live Dataview tables: warm/dormant ties, network-by-company, a by-source breakdown._DATA_POINTS.md— every node type + relation + a field-enrichment-by-source matrix + a "mineable questions" catalog._GRAPH.md+--graph-config— colors the global graph by source and type (merges into.obsidian/graph.json, preserving your settings).copilot-prompts/— one-click/commandsfor Obsidian Copilot:/warm-intro,/investor-paths,/reconnect,/job-fit,/mine,/for-me.
Obsidian integration
Two ways to put an AI directly in your vault — both install from Settings → Community plugins:
Semantic Vault QA across all notes, plus the generated /commands. Best for instant retrieval and one-shot questions.
Claude Code itself in a side panel, with your brain as its working directory. Best for multi-step reasoning, drafting, and writing back into the vault.
The answer can be an artifact, not just a chat reply. Because the agent can write into the vault, ask for the output as a file: a .canvas board of your warmest investor paths, a live Dataview dashboard of dormant contacts, a note comparing 50-mirror/ to how you describe yourself. Whatever it generates follows the builder's conventions, so it's instantly first-class in the graph.
One honest limit: the export is a snapshot — sharp on your history, blind to this week. Re-export periodically and update in place with --refresh (or Studio's Update reseed) — your edits are kept. Hosted, scheduled always-fresh sync is on the roadmap.
CLI reference — under the hood
You normally never type these — you prompt the skill and it runs the engine for you (and adapts the commands to your export). They're documented for transparency, for self-hosters, and for the AI itself: the engine is plain Python and every command also runs without the Agent Skill installed.
Flags
| Flag | What it does |
|---|---|
--full | Owner mode — keep your own emails/phones and quarantined personal files (still 100% local) |
--emit obsidian|gbrain|both | Output target — Obsidian vault (default) and/or a GBrain repo |
--subject person|company | What the brain is rooted on (00-me/ vs 00-org/) |
--provider claude|openai | Which in-vault guide to write (CLAUDE.md vs AGENTS.md) |
--mappings <dir> | Override or add declarative JSON source mappings |
--dry-run | Preview detection without writing |
--doctor | Preflight check |
Each output dir must be new or empty — the builder never clobbers existing notes.
Adding a source
Anything with a data export can become part of the brain. In order of preference:
- A declarative JSON mapping —
engine/mappings/sources/<name>.json: adetectblock plusrecords[]thatlocatearrays and pluckfieldsvia a tiny selector language into canonicalcol.add_*verbs. No Python, no engine edit. A mapping wins over a same-named Python adapter and can ship via--mappings <dir>. - A Python adapter — only for bespoke cross-file logic: one file in
engine/scripts/sources/personal/or…/company/(setSUBJECT="company"for company sources). Auto-discovered, no registry edit. - An output target — one emitter file in
engine/scripts/emitters/.
Either way the builder, privacy rules, Obsidian/GBrain output, and cross-source merging all work unchanged. Files no mapping claims are caught by the universal harvester and listed in _COVERAGE.md under "Needs a mapping" — your cue to add a rule for precision.
Project layout
One provider-neutral engine does the work; thin per-provider manifests ship it as a cross-model Agent Skill (the Agent Skills open standard).
Privacy model
- Zero network calls in the core transform — verifiable in
engine/scripts/. - Third-party PII stripped at parse time — across every source, contacts' emails and phones never become notes.
- Message bodies never written — only a per-person count + last-contact date, which grades relationship strength.
- Sensitive files never imported —
Logins,Receipts,Security Challenges,ImportedContactsand friends stay in your original archive, listed in_quarantine/. - The schema map redacts sensitive columns and masks any email/phone in samples.
Never commit your vault to a public repo — the generated .gitignore guards against accidents. Full write-up: the privacy model.
Roadmap
LinkedIn → vault: full local transform, Obsidian-native
shippedMulti-source: Facebook, Instagram, Google — one merged vault
shippedPersonal and company; GBrain emitter; --full owner mode
Cross-model (Claude + Codex); multi-entity + _correlations/; JSON mappings + harvester
24 sources: X, WhatsApp, GitHub, YouTube, Strava, Reddit, Spotify, TikTok + Notion, Confluence, Jira, Salesforce, HubSpot, Zendesk, Email, Microsoft 365, Teams; offline geocoding; company-named vault layout; --refresh re-import
Always fresh: --refresh re-import shipped — update the brain in place from a newer archive, your edits kept (hosted scheduled sync lands with Cloud)
Sharper entity resolution: stable IDs + precision-biased fuzzy matching
plannedThe agentic ops layer — the brain that acts: meeting prep, drafting in your voice, reviving dormant relationships
nextContributing
The most valuable contribution right now: run it on your real export and open an issue if any file or column didn't map cleanly — the schema_map.md it generates is exactly what we need to see. Parser robustness across the long tail of real accounts is how this gets great.
Created and maintained by Adrian Vicovan. Questions? Use our contact form.