The LessWrong Archive

Every post and comment on LessWrong, from the 2003 essays later imported from Overcoming Bias, through LessWrong 1.0, to a quarter of an hour ago. Pulled from the site's public API with every field it exposes: bodies as HTML and Markdown, karma, vote counts and reacts, tags, curation and review state, Alignment Forum flags, the legacy identifiers. One zstd-compressed ndjson file per month, posts and comments together, held at the Laboratory for Social Minds, CMU.

CMU

The files are served only to Carnegie Mellon IP ranges — on-campus networks and the CMU VPN; from anywhere else they return 403. The index this page reads is public. The writing belongs to its authors and was published on LessWrong under that site's terms; this copy is held for research at CMU, not redistributed.

connecting to the archive index…

The archive, by year

One link per month; hover for counts and size. Dashed months have nothing in them — before 2009 the site holds only the essays it later imported. Amber months are still being fetched, and the green-edged month is the current one, which grows through the month. The grid is drawn from the archive index and redraws itself every minute.

The whole archiveloading…
Each record is fetched once, and once more if it was young. Everything in the backfill was already settled when it was taken, and is stored as it stood that day. New posts and comments are picked up within fifteen minutes of appearing, when they have few votes; thirty days after posting each is fetched a second and final time, so the archive ends up holding mature karma, reacts and text for everything. Nothing is requested a third time, and edits or deletions after that point are not reflected. fetched_at on each record says when its copy was taken; the index lists each file's SHA-256.

Using the archive

On ganesha (kubrick, a 20 TB disk at /mnt/kubrick)

PathContents
/mnt/kubrick/lesswrong/dumps/LW_2003-01 … the current month, LW_YYYY-MM.jsonl.zst
/mnt/kubrick/lesswrong/meta/tags, sequences, collections (the editorial structure), refreshed weekly
/mnt/kubrick/lesswrong/lesswrong.sqliteevery record, latest version, indexed by month, post and user — the source the files are written from
/mnt/kubrick/lesswrong/index.jsonthe month list with counts, sizes and hashes, plus the puller's status
/mnt/kubrick/README.mdhow the archive is held, served and extended

Reading a dump

zstd -dc LW_2012-03.jsonl.zst | head -1 | python3 -m json.tool | head -40
zstd -dc LW_2012-03.jsonl.zst | jq -c 'select(.type=="post") | {title, baseScore, voteCount, user: .user.username}'

Plain zstd -dc works; no long window is needed. Each line is one JSON object: a type of post or comment, a fetched_at timestamp, then the record exactly as the LessWrong API returned it. Lines are ordered by postedAt, so a month reads chronologically. Comments carry postId, parentCommentId and topLevelCommentId; a thread whose post and comments fall in one month rebuilds from that file alone, and the SQLite joins across months for the rest.

What a record holds

GroupFields (posts)Fields (comments)
Identity_id slug title userId user{username, displayName, karma, createdAt} coauthors legacy legacyId legacyData_id postId parentCommentId topLevelCommentId userId user{…} legacyId legacyParentId
TimepostedAt modifiedAt createdAt lastCommentedAt curatedDate frontpageDate scoreExceeded…DatepostedAt lastEditedAt lastSubthreadActivity promotedAt deletedDate
Textcontents{html, markdown, wordCount, version, editedAt} htmlBody tableOfContentscontents{html, markdown, wordCount, version, editedAt} htmlBody
ReceptionbaseScore voteCount score extendedScore (reacts, agreement) maxBaseScore commentCount viewCount afBaseScore afVoteCountbaseScore voteCount score extendedScore directChildrenCount descendentCount afBaseScore
Editorialtags{name, slug, core} tagRelevance frontpageDate curatedDate reviewCount reviewVoteScore… finalReviewVote… reviewWinner spotlight canonicalSequenceId canonicalCollectionSlug sticky meta question shortform isEvent afanswer promoted moderatorHat shortform af relevantTagIds nominatedForReview reviewingForReview
Moderationstatus draft unlisted rejected rejectedReason bannedUserIds commentsLocked authorIsUnrevieweddeleted deletedPublic deletedReason retracted spam rejected needsReview repliesBlockedUntil

A field that does not apply is present as null, so readers can rely on the keys. Not here, because the API does not expose it to an anonymous reader: individual votes (only their aggregates), drafts, private messages, and anything deleted before the 2017 move to the current site.

Source and method

LessWrong's public GraphQL endpoint (https://www.lesswrong.com/graphql), read one request per second with an identifying User-Agent by the lesswrong-pull service on ganesha. Posts come from the site's new view and comments from allRecentComments, each walked month by month and page by page on postedAt; posts the lists leave out but comments refer to (shortform containers, unlisted posts) are fetched by id. The backfill ran oldest month first. Since then the puller asks every fifteen minutes for anything newer than the newest record it holds, and once a day re-fetches the records that have just turned thirty days old. New records go into SQLite and the month they belong to is rewritten from it. LessWrong shares its software with the EA Forum and the Alignment Forum, so the same tool would cover either.

Documentation