Data Efficacy for Language Model Training
-
Updated
Oct 10, 2026 - Python
Data Efficacy for Language Model Training
Curation of BIDS (CuBIDS): A sanity-preserving software package for processing BIDS datasets.
Application to help you track, categorize, and rate the fanmade music videos you watch (with a focus on anime music videos)
A file management system for turning digital chaos into an organized archive. Includes scripts to sort, analyze, and consolidate files.
Sanitize and organize astrophotography subframes from ASIAir or DSLR captured data. Organize files by folder for WBPP keywords, cleanup empty directories, remove jpg previews, rename RAW files using EXIF data.
This application is designed to simplify the process of collecting and managing leads for an organization. It provides an intuitive user interface and several useful features to streamline data entry, organization, and follow-up activities.
A React + Vite application for managing hospital appointments. Add patients, view patient lists, and navigate between pages seamlessly.
Organize experimental data in a structured and coherent manner
Attempts to correct the arbitrary filenames of photos/videos/audio downloaded from Instagram based on the time they were sent.
copyright-stats-extractor parses headlines/articles on digital copyright enforcement to auto‑extract stats like takedown counts, year, and parties.
Extracts key release details from unstructured text to create clear, structured summaries.
Applying data visualization techniques; organizing, cleaning, and aggregating data from multiple sources; I determine how much happiness has changed over time in all regions of the world
rankextractplus extracts and structures ranked info from text, organizing data for easier comparison and analysis.
Echo is a versatile information recording and management software that empowers users to efficiently organize and access their data.
15 hand postures from 8-channel surface EMG. I audited my own evaluation and found the split was leaking overlapping windows between train and test — the honest number is 14.5 points below what I first reported. 83.7% within-subject, 32.3% cross-subject.
Access computer science history by year, including major breakthroughs, research papers, and technological advancements.
Python automation scripts: web scraping for data analysis
To associate your repository with the data-organization topic, visit your repo's landing page and select "manage topics."