In-depth AI/ML paper reviews, summaries, and tech guides β published as a blog.
A Jekyll static site, deployed to GitHub Pages.
π°π· νκ΅μ΄ README
If you've never run a Jekyll site before, here's the whole loop.
1. Install Ruby 3.3+ and Bundler. Check what you have:
ruby --version # need 3.3 or newer (see .ruby-version)
bundle --version # ships with Ruby; if missing: gem install bundlerOn macOS the system Ruby is old β use rbenv or
asdf to install 3.3+. (The repo pins the version in .ruby-version, so a version
manager will pick it up automatically.)
2. Install the project's gems (Jekyll, plugins, html-proofer):
bundle install3. Run the dev server. It rebuilds on save and serves at
http://localhost:4000:
bundle exec jekyll serveEdit a file under _posts/, _sass/, or _includes/, save, and refresh β most
changes appear immediately. (Changes to _config.yml need a server restart.)
4. Build for production (what CI does) when you want the final output in
_site/:
bundle exec jekyll buildWhy plain
jekylland notgithub-pages? This site uses custom Ruby plugins in_plugins/, which the sandboxedgithub-pagesgem disallows. So both local builds and CI run Jekyll directly.
_posts/ Posts β YYYY-MM-DD-slug.md (Korean; English twin is -en.md)
_layouts/ Page templates: default β post / page
_includes/ Reusable fragments: head, header, footer, nav_links,
page_divider, category-posts, language_switcher,
related_posts, icon (inline SVG icons)
_sass/ Styles: _layout, _post, _tags, _syntax (Rouge code theme),
_dark (dark mode), base/*
β bourbon/ and neat/ are vendored frameworks β don't edit
_plugins/ reading_time.rb (KO/EN-aware read time)
lazy_images.rb (adds loading="lazy" to <img>)
image_dimensions.rb (width/height on local <img>)
post_description.rb (fills page.description for posts)
related_posts.rb (page.related, and the prev/next links)
scrollable_tables.rb (wraps wide tables so they scroll)
search_index.rb (plain_text filter for search.json)
legacy_urls.rb (redirects pre-slug post URLs)
translations.rb (keeps a post's translation out of the lists)
css/ main.scss (Sass entry point) Β· search.css (search page only)
js/ main.js (theme toggle, code-copy, TOC, menu, image zoomβ¦)
search.js (drives the search box)
assets/images/ Topic cover images, plus post figures and diagrams (flat)
assets/<slug>/ Figures of some older posts; new posts use assets/images/
search.json Full-text search index (consumed by simple-jekyll-search)
test/ minitest unit tests for the _plugins/ logic
script/ lint-posts.rb (post sources vs the conventions, before the build)
validate-site.sh (post-build discoverability checks)
_data/ post_conventions.yml (categories, covers, controlled and
retired tags β the one copy, also read by scholar-lens)
.github/workflows/ CI: tests β lint β build β html-proofer β validate-site; deploys on push to main
Top-level pages: index.html (home), plus paper-reviews.md,
paper-summaries.md, tech-guides.md, insights.md (the four section pages),
categories.html, tags.html, search.md, about.md, and 404.html.
sitemap.xml, feed.xml and robots.txt are generated by the plugins; there is
no source file for any of them.
The easiest path is the /write-post skill, which runs the whole
research β draft β proofread workflow. To add one by hand, create
_posts/YYYY-MM-DD-slug.md starting with this front matter:
---
layout: post
title: "<Post Title>"
subtitle: "<one-line pitch>" # optional β shown under the title in the header
date: YYYY-MM-DD HH:MM:SS
paper_author: "<Author>" # the paper's org; omit for Insights/opinion posts
seo_title: "<shorter title>" # optional β for <title> when the paper title is too long
description: >- # optional β see below
<search-snippet, ~150 chars>
categories: ["<Type>", "<Topic>"]
tags: ["<Tag-1>", "<Tag-2>"]
cover: /assets/images/<topic>.(jpg|png)
use_math: true # ONLY if the post has equations (loads MathJax)
lang: ko # required with translation_id (ko | en)β¦
translation_id: <shared-slug> # β¦links a Korean post to its -en twin;
# only the Korean one is listed
---A paper post's <title> is its title (or seo_title) plus "λ
Όλ¬Έ 리뷰" / "λ
Όλ¬Έ
μμ½", the words a reader searches with; social cards keep the plain title.
Don't repeat the title as an H1 in the body. The layout already renders it,
so a leading # Title produces two <h1>s and leaks into the search snippet.
Put a tagline in subtitle: instead.
description: is what Google shows under the link, what social cards quote, and
what the RSS <summary> carries. If you omit it, _plugins/post_description.rb
derives one from the post's first real prose paragraph, which is usually good
enough. Write it by hand when the first paragraph opens on a pull quote or a
disclosure note β that is, on most Insights posts.
Descriptions must be unique across the site; CI fails the build if two pages share one.
Categories are two levels:
categories[0]β the type:Paper Reviews,Paper Summaries,Tech Guides, orInsights. This decides which nav tab the post appears under. A type with no posts yet keeps its tab and renders an empty-state line: a missing tab reads as a section that was removed, not one still filling up.categories[1]β the topic:Language-Models,Multimodal-Learning,Finetuning,Retrieval-Augmented-Generation,Agentic-AI,Data-Architecture, β¦ (add new ones freely).
Jekyll slugifies the two and combines them with the date to build the output
path (permalink in _config.yml):
categories: ["Paper Reviews", "Language-Models"] + date: 2025-01-23
β
_site/paper-reviews/language-models/2025/01/23/<slug>.html
So changing the categories or date of a published post changes its URL, which breaks inbound links and search results. Set them once and leave them.
Posts first published under the old unslugified form (/paper%20reviews/β¦) keep
working: _plugins/legacy_urls.rb adds that path to each post's redirect_from,
and jekyll-redirect-from writes a stub there that redirects to the new URL.
Tags are free-form and hyphenated, and a tag phrased as one paper's contribution
(Fine-Grained-Expert-Segmentation) can only ever apply to that paper. Those
make a precise index and connect nothing, and _plugins/related_posts.rb
requires a shared tag β so a post tagged only that way ships with no
"Related reading" block at all.
So also give each post at least one tag from the controlled topic layer,
tags.controlled in _data/post_conventions.yml.
That file is the one copy of the list; script/lint-posts.rb and scholar-lens
both read it.
Keep the specific tags β they say something the topic tag does not. Add the topic tag, don't swap for it.
Before coining a new tag, check that an existing one doesn't already cover it.
Three names for one idea (Agentic-Architecture, Agentic-Patterns,
Agentic-Infrastructure) leave every post holding a tag no other post shares,
which is the same as having no topic tag at all. When two tags turn out to name
one idea, keep one and add the other to tags.aliases; the lint then rejects the
retired name.
Write $$β¦$$ for both inline and display math, and set use_math: true.
Never use single $β¦$. kramdown doesn't treat single $ as math, so its
Markdown pass turns _/* inside the span into <em>/<strong> before
MathJax runs β e.g. $a*b*c$ becomes $a<em>b</em>c$ and renders broken. With
$$, kramdown emits verbatim \(β¦\) and leaves the contents alone. (Prose
dollar signs like $10M are fine β they're not math.)
These are the gates CI runs, in the same order. Build to a throwaway
directory rather than _site/, for the reason in the note below:
ruby test/run_all.rb # plugin logic still correct?
ruby script/lint-posts.rb # post sources follow the conventions?
bundle exec jekyll build --strict-front-matter \
--destination /tmp/site-verify # does it build clean?
bundle exec htmlproofer /tmp/site-verify --disable-external \
--allow-hash-href --no-enforce-https \
--ignore-urls "/^\/page\/\d+/" # broken links, images, anchors?
script/validate-site.sh /tmp/site-verify # sitemap, feed, metadata, headingsIf a check reports something impossible, look for a running
jekyll servefirst. It watches the tree and rewrites_site/behind you, it overridessite.urlwithhttp://localhost:4000(so every sitemap URL looks wrong), and it holds the_config.ymlit started with β soexcludeentries added since then don't apply. Building elsewhere sidesteps all three:ps aux | grep '[j]ekyll serve'
test/ unit-tests the pure logic in _plugins/ β one file per plugin. Plain
ruby, not bundle exec: each plugin guards its Jekyll/Liquid registration
behind defined? so the logic loads standalone, and minitest ships with Ruby.
Anything you change in _plugins/ changes every page on the site, so add a
case before changing behaviour.
.github/workflows/jekyll.yml runs on pull requests to main as well as
pushes to it, so the gates below block a bad merge rather than merely reporting
one after the fact. Steps 1β5 run on both events; step 6 is skipped for pull
requests. In order, it:
- runs
ruby test/run_all.rb(the_plugins/unit tests), - runs
ruby script/lint-posts.rbβ front matter against_data/post_conventions.yml, plus table, math and heading rules in the Markdown source, - builds the site with
JEKYLL_ENV=production, - runs html-proofer over
_site/(internal links, images, anchors), - runs
script/validate-site.shβ sitemap/feed parse at byte 0, every sitemap URL under the configuredurl, a pinned build timezone,robots.txtnot blocking and naming the sitemap, at least one rendered page, exactly oneh1per page, no heading-level skips, a description and a canonical on every page, every description longer than its own title, no duplicate description or title, no listing page showing a translation beside its original, and no authoring sources published β and - deploys to GitHub Pages.
A failure is almost always step 4 or 5; the Actions log names the exact link, image, or page. There is no manual deploy step.
β Don't add
google*.html/naver*.htmlto_config.yml'sexclude. They're Search Console / Naver ownership-verification tokens that must ship to the site root. Excluding them silently breaks ownership verification.
Check the file first β it is usually fine:
curl -sI https://bits-bytes-nn.github.io/sitemap.xml # expect 200, application/xml
curl -sS https://bits-bytes-nn.github.io/sitemap.xml -o /tmp/s.xml && \
ruby -rrexml/document -e 'REXML::Document.new(File.read("/tmp/s.xml")); puts "well-formed"'
curl -sS https://bits-bytes-nn.github.io/robots.txtIf those pass, the site is not the problem, and no change to the file will fix
it. "Couldn't fetch" is Search Console reporting that Google has not fetched the
sitemap. That is a crawl-scheduling decision on Google's side, not a failed
request. It is common on *.github.io properties, and re-submitting under
another URL does not change it.
What to do instead:
- Submit
https://bits-bytes-nn.github.io/sitemap.xmlonce, and leave it. Delete any other sitemap entries. - Judge indexing by Indexing β Pages and by URL Inspection on a few
posts, not by the sitemap row. Posts are reachable from the home page, the
section pages and
/tags/, androbots.txtnames the sitemap, so Google can find every post without a processed sitemap. - The only change reported to make the sitemap fetch is moving the site to a custom domain verified as a Domain property. That changes every URL, so it needs redirects and is a decision of its own.
MIT β see LICENSE. Built on the Centrarium Jekyll theme.
