Skip to content

About

In-depth AI/ML paper reviews, tech guides, and insights on the latest research and engineering.

Topics

Resources

Stars

1 star

Watchers

1 watching

Forks

Latest commit

Β 

History

278 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

🧠 Bits, Bytes and Neural Networks

In-depth AI/ML paper reviews, summaries, and tech guides β€” published as a blog.

A Jekyll static site, deployed to GitHub Pages.

Deploy Ruby Jekyll GitHub Pages

πŸ‡°πŸ‡· ν•œκ΅­μ–΄ README

Bits, Bytes and Neural Networks


Quick start

If you've never run a Jekyll site before, here's the whole loop.

1. Install Ruby 3.3+ and Bundler. Check what you have:

ruby --version     # need 3.3 or newer (see .ruby-version)
bundle --version   # ships with Ruby; if missing: gem install bundler

On macOS the system Ruby is old β€” use rbenv or asdf to install 3.3+. (The repo pins the version in .ruby-version, so a version manager will pick it up automatically.)

2. Install the project's gems (Jekyll, plugins, html-proofer):

bundle install

3. Run the dev server. It rebuilds on save and serves at http://localhost:4000:

bundle exec jekyll serve

Edit a file under _posts/, _sass/, or _includes/, save, and refresh β€” most changes appear immediately. (Changes to _config.yml need a server restart.)

4. Build for production (what CI does) when you want the final output in _site/:

bundle exec jekyll build

Why plain jekyll and not github-pages? This site uses custom Ruby plugins in _plugins/, which the sandboxed github-pages gem disallows. So both local builds and CI run Jekyll directly.


Project structure

_posts/            Posts β€” YYYY-MM-DD-slug.md (Korean; English twin is -en.md)
_layouts/          Page templates: default β†’ post / page
_includes/         Reusable fragments: head, header, footer, nav_links,
                   page_divider, category-posts, language_switcher,
                   related_posts, icon (inline SVG icons)
_sass/             Styles: _layout, _post, _tags, _syntax (Rouge code theme),
                   _dark (dark mode), base/*
                   ⚠ bourbon/ and neat/ are vendored frameworks β€” don't edit
_plugins/          reading_time.rb      (KO/EN-aware read time)
                   lazy_images.rb       (adds loading="lazy" to <img>)
                   image_dimensions.rb  (width/height on local <img>)
                   post_description.rb  (fills page.description for posts)
                   related_posts.rb     (page.related, and the prev/next links)
                   scrollable_tables.rb (wraps wide tables so they scroll)
                   search_index.rb      (plain_text filter for search.json)
                   legacy_urls.rb       (redirects pre-slug post URLs)
                   translations.rb      (keeps a post's translation out of the lists)
css/               main.scss (Sass entry point) Β· search.css (search page only)
js/                main.js (theme toggle, code-copy, TOC, menu, image zoom…)
                   search.js (drives the search box)
assets/images/     Topic cover images, plus post figures and diagrams (flat)
assets/<slug>/     Figures of some older posts; new posts use assets/images/
search.json        Full-text search index (consumed by simple-jekyll-search)
test/              minitest unit tests for the _plugins/ logic
script/            lint-posts.rb (post sources vs the conventions, before the build)
                   validate-site.sh (post-build discoverability checks)
_data/             post_conventions.yml (categories, covers, controlled and
                   retired tags β€” the one copy, also read by scholar-lens)
.github/workflows/ CI: tests β†’ lint β†’ build β†’ html-proofer β†’ validate-site; deploys on push to main

Top-level pages: index.html (home), plus paper-reviews.md, paper-summaries.md, tech-guides.md, insights.md (the four section pages), categories.html, tags.html, search.md, about.md, and 404.html. sitemap.xml, feed.xml and robots.txt are generated by the plugins; there is no source file for any of them.


Writing a post

The easiest path is the /write-post skill, which runs the whole research β†’ draft β†’ proofread workflow. To add one by hand, create _posts/YYYY-MM-DD-slug.md starting with this front matter:

---
layout: post
title: "<Post Title>"
subtitle: "<one-line pitch>"       # optional β€” shown under the title in the header
date: YYYY-MM-DD HH:MM:SS
paper_author: "<Author>"           # the paper's org; omit for Insights/opinion posts
seo_title: "<shorter title>"       # optional β€” for <title> when the paper title is too long
description: >-                    # optional β€” see below
  <search-snippet, ~150 chars>
categories: ["<Type>", "<Topic>"]
tags: ["<Tag-1>", "<Tag-2>"]
cover: /assets/images/<topic>.(jpg|png)
use_math: true                     # ONLY if the post has equations (loads MathJax)
lang: ko                           # required with translation_id (ko | en)…
translation_id: <shared-slug>      # …links a Korean post to its -en twin;
                                   #    only the Korean one is listed
---

A paper post's <title> is its title (or seo_title) plus "λ…Όλ¬Έ 리뷰" / "λ…Όλ¬Έ μš”μ•½", the words a reader searches with; social cards keep the plain title.

Don't repeat the title as an H1 in the body. The layout already renders it, so a leading # Title produces two <h1>s and leaks into the search snippet. Put a tagline in subtitle: instead.

description: β€” the search snippet

description: is what Google shows under the link, what social cards quote, and what the RSS <summary> carries. If you omit it, _plugins/post_description.rb derives one from the post's first real prose paragraph, which is usually good enough. Write it by hand when the first paragraph opens on a pull quote or a disclosure note β€” that is, on most Insights posts.

Descriptions must be unique across the site; CI fails the build if two pages share one.

Categories drive the URL

Categories are two levels:

  • categories[0] β€” the type: Paper Reviews, Paper Summaries, Tech Guides, or Insights. This decides which nav tab the post appears under. A type with no posts yet keeps its tab and renders an empty-state line: a missing tab reads as a section that was removed, not one still filling up.
  • categories[1] β€” the topic: Language-Models, Multimodal-Learning, Finetuning, Retrieval-Augmented-Generation, Agentic-AI, Data-Architecture, … (add new ones freely).

Jekyll slugifies the two and combines them with the date to build the output path (permalink in _config.yml):

categories: ["Paper Reviews", "Language-Models"] + date: 2025-01-23
        ↓
_site/paper-reviews/language-models/2025/01/23/<slug>.html

So changing the categories or date of a published post changes its URL, which breaks inbound links and search results. Set them once and leave them.

Posts first published under the old unslugified form (/paper%20reviews/…) keep working: _plugins/legacy_urls.rb adds that path to each post's redirect_from, and jekyll-redirect-from writes a stub there that redirects to the new URL.

Tags: one topic tag on top of the specific ones

Tags are free-form and hyphenated, and a tag phrased as one paper's contribution (Fine-Grained-Expert-Segmentation) can only ever apply to that paper. Those make a precise index and connect nothing, and _plugins/related_posts.rb requires a shared tag β€” so a post tagged only that way ships with no "Related reading" block at all.

So also give each post at least one tag from the controlled topic layer, tags.controlled in _data/post_conventions.yml. That file is the one copy of the list; script/lint-posts.rb and scholar-lens both read it.

Keep the specific tags β€” they say something the topic tag does not. Add the topic tag, don't swap for it.

Before coining a new tag, check that an existing one doesn't already cover it. Three names for one idea (Agentic-Architecture, Agentic-Patterns, Agentic-Infrastructure) leave every post holding a tag no other post shares, which is the same as having no topic tag at all. When two tags turn out to name one idea, keep one and add the other to tags.aliases; the lint then rejects the retired name.

Math: always use $$…$$

Write $$…$$ for both inline and display math, and set use_math: true.

Never use single $…$. kramdown doesn't treat single $ as math, so its Markdown pass turns _/* inside the span into <em>/<strong> before MathJax runs β€” e.g. $a*b*c$ becomes $a<em>b</em>c$ and renders broken. With $$, kramdown emits verbatim \(…\) and leaves the contents alone. (Prose dollar signs like $10M are fine β€” they're not math.)

Validate before pushing

These are the gates CI runs, in the same order. Build to a throwaway directory rather than _site/, for the reason in the note below:

ruby test/run_all.rb                                        # plugin logic still correct?
ruby script/lint-posts.rb                                   # post sources follow the conventions?
bundle exec jekyll build --strict-front-matter \
  --destination /tmp/site-verify                            # does it build clean?
bundle exec htmlproofer /tmp/site-verify --disable-external \
  --allow-hash-href --no-enforce-https \
  --ignore-urls "/^\/page\/\d+/"                            # broken links, images, anchors?
script/validate-site.sh /tmp/site-verify                    # sitemap, feed, metadata, headings

If a check reports something impossible, look for a running jekyll serve first. It watches the tree and rewrites _site/ behind you, it overrides site.url with http://localhost:4000 (so every sitemap URL looks wrong), and it holds the _config.yml it started with β€” so exclude entries added since then don't apply. Building elsewhere sidesteps all three:

ps aux | grep '[j]ekyll serve'

test/ unit-tests the pure logic in _plugins/ β€” one file per plugin. Plain ruby, not bundle exec: each plugin guards its Jekyll/Liquid registration behind defined? so the logic loads standalone, and minitest ships with Ruby. Anything you change in _plugins/ changes every page on the site, so add a case before changing behaviour.


Deployment

.github/workflows/jekyll.yml runs on pull requests to main as well as pushes to it, so the gates below block a bad merge rather than merely reporting one after the fact. Steps 1–5 run on both events; step 6 is skipped for pull requests. In order, it:

  1. runs ruby test/run_all.rb (the _plugins/ unit tests),
  2. runs ruby script/lint-posts.rb β€” front matter against _data/post_conventions.yml, plus table, math and heading rules in the Markdown source,
  3. builds the site with JEKYLL_ENV=production,
  4. runs html-proofer over _site/ (internal links, images, anchors),
  5. runs script/validate-site.sh β€” sitemap/feed parse at byte 0, every sitemap URL under the configured url, a pinned build timezone, robots.txt not blocking and naming the sitemap, at least one rendered page, exactly one h1 per page, no heading-level skips, a description and a canonical on every page, every description longer than its own title, no duplicate description or title, no listing page showing a translation beside its original, and no authoring sources published β€” and
  6. deploys to GitHub Pages.

A failure is almost always step 4 or 5; the Actions log names the exact link, image, or page. There is no manual deploy step.

⚠ Don't add google*.html / naver*.html to _config.yml's exclude. They're Search Console / Naver ownership-verification tokens that must ship to the site root. Excluding them silently breaks ownership verification.

If Search Console says it can't fetch the sitemap

Check the file first β€” it is usually fine:

curl -sI  https://bits-bytes-nn.github.io/sitemap.xml   # expect 200, application/xml
curl -sS  https://bits-bytes-nn.github.io/sitemap.xml -o /tmp/s.xml && \
  ruby -rrexml/document -e 'REXML::Document.new(File.read("/tmp/s.xml")); puts "well-formed"'
curl -sS  https://bits-bytes-nn.github.io/robots.txt

If those pass, the site is not the problem, and no change to the file will fix it. "Couldn't fetch" is Search Console reporting that Google has not fetched the sitemap. That is a crawl-scheduling decision on Google's side, not a failed request. It is common on *.github.io properties, and re-submitting under another URL does not change it.

What to do instead:

  • Submit https://bits-bytes-nn.github.io/sitemap.xml once, and leave it. Delete any other sitemap entries.
  • Judge indexing by Indexing β†’ Pages and by URL Inspection on a few posts, not by the sitemap row. Posts are reachable from the home page, the section pages and /tags/, and robots.txt names the sitemap, so Google can find every post without a processed sitemap.
  • The only change reported to make the sitemap fetch is moving the site to a custom domain verified as a Domain property. That changes every URL, so it needs redirects and is a decision of its own.

License

MIT β€” see LICENSE. Built on the Centrarium Jekyll theme.

About

In-depth AI/ML paper reviews, tech guides, and insights on the latest research and engineering.

Topics

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages