Every post and comment on LessWrong, from the 2003 essays later imported from Overcoming Bias, through LessWrong 1.0, to a quarter of an hour ago. Pulled from the site's public API with every field it exposes: bodies as HTML and Markdown, karma, vote counts and reacts, tags, curation and review state, Alignment Forum flags, the legacy identifiers. One zstd-compressed ndjson file per month, posts and comments together, held at the Laboratory for Social Minds, CMU.
The files are served only to Carnegie Mellon IP ranges — on-campus networks and the
CMU VPN; from anywhere else they return 403. The index this page reads is public. The
writing belongs to its authors and was published on LessWrong under that site's terms; this copy is
held for research at CMU, not redistributed.
One link per month; hover for counts and size. Dashed months have nothing in them — before 2009 the site holds only the essays it later imported. Amber months are still being fetched, and the green-edged month is the current one, which grows through the month. The grid is drawn from the archive index and redraws itself every minute.
fetched_at on each record says when its copy was
taken; the index lists each file's SHA-256./mnt/kubrick)| Path | Contents |
|---|---|
/mnt/kubrick/lesswrong/dumps/ | LW_2003-01 … the current month, LW_YYYY-MM.jsonl.zst |
/mnt/kubrick/lesswrong/meta/ | tags, sequences, collections (the editorial structure), refreshed weekly |
/mnt/kubrick/lesswrong/lesswrong.sqlite | every record, latest version, indexed by month, post and user — the source the files are written from |
/mnt/kubrick/lesswrong/index.json | the month list with counts, sizes and hashes, plus the puller's status |
/mnt/kubrick/README.md | how the archive is held, served and extended |
zstd -dc LW_2012-03.jsonl.zst | head -1 | python3 -m json.tool | head -40
zstd -dc LW_2012-03.jsonl.zst | jq -c 'select(.type=="post") | {title, baseScore, voteCount, user: .user.username}'
Plain zstd -dc works; no long
window is needed. Each line is one JSON object: a type of post or
comment, a fetched_at timestamp, then the record exactly as the LessWrong
API returned it. Lines are ordered by postedAt, so a month reads chronologically.
Comments carry postId, parentCommentId and topLevelCommentId;
a thread whose post and comments fall in one month rebuilds from that file alone, and the SQLite
joins across months for the rest.
| Group | Fields (posts) | Fields (comments) |
|---|---|---|
| Identity | _id slug title userId user{username, displayName, karma, createdAt} coauthors legacy legacyId legacyData | _id postId parentCommentId topLevelCommentId userId user{…} legacyId legacyParentId |
| Time | postedAt modifiedAt createdAt lastCommentedAt curatedDate frontpageDate scoreExceeded…Date | postedAt lastEditedAt lastSubthreadActivity promotedAt deletedDate |
| Text | contents{html, markdown, wordCount, version, editedAt} htmlBody tableOfContents | contents{html, markdown, wordCount, version, editedAt} htmlBody |
| Reception | baseScore voteCount score extendedScore (reacts, agreement) maxBaseScore commentCount viewCount afBaseScore afVoteCount | baseScore voteCount score extendedScore directChildrenCount descendentCount afBaseScore |
| Editorial | tags{name, slug, core} tagRelevance frontpageDate curatedDate reviewCount reviewVoteScore… finalReviewVote… reviewWinner spotlight canonicalSequenceId canonicalCollectionSlug sticky meta question shortform isEvent af | answer promoted moderatorHat shortform af relevantTagIds nominatedForReview reviewingForReview |
| Moderation | status draft unlisted rejected rejectedReason bannedUserIds commentsLocked authorIsUnreviewed | deleted deletedPublic deletedReason retracted spam rejected needsReview repliesBlockedUntil |
A field that does not apply is present
as null, so readers can rely on the keys. Not here, because the API does not expose it
to an anonymous reader: individual votes (only their aggregates), drafts, private messages, and
anything deleted before the 2017 move to the current site.
LessWrong's public GraphQL endpoint
(https://www.lesswrong.com/graphql), read one request per second with an identifying
User-Agent by the lesswrong-pull service on ganesha. Posts come from the site's
new view and comments from allRecentComments, each walked month by month
and page by page on postedAt; posts the lists leave out but comments refer to
(shortform containers, unlisted posts) are fetched by id. The backfill ran oldest month first.
Since then the puller asks every fifteen minutes for anything newer than the newest record it
holds, and once a day re-fetches the records that have just turned thirty days old. New records go
into SQLite and the month they belong to is rewritten from it. LessWrong shares its software with
the EA Forum and the Alignment Forum, so the same tool would cover either.