Wes Ellis./ a personal notebook
Technology. Stories. Side projects.
A few things worth writing down.
← Back to Side Projects

Side Projects

FlowDrive: A Storage Optimizer That Only Exists on Paper (So Far)

A disassembled hard drive with its platter exposed next to a circuit board and screwdriver on a workbench.

Part 1 of the thread App experiments

PROJECT AT A GLANCEOpen source · concept
What it is
Desktop app (concept)
My role
Creator
Year
2026
Built with
  • Python (planned)
  • PyQt6 (planned)
  • scikit-learn (planned)
  • C++ or Rust (undecided)
THE SHORT VERSION4 points
  • FlowDrive is an idea for a tool that predicts which files you'll open and moves them to faster storage before you need them.
  • Right now the repo is documentation only: a spec, a technical appendix, an architecture doc and a roadmap. No code.
  • The plan is a C++ or Rust core engine, a Python ML layer and a PyQt6 interface. The C++ versus Rust call is still open.
  • Treat the feature list as design goals, not things you can install.

If you've got a small fast SSD and a big slow hard drive, you already know the chore. Something you use every day ends up on the slow drive, something you haven't touched in a year is hogging the fast one, and you're the one who has to notice and shuffle things around. I wanted to write down what it would look like if the computer just figured that out.

flowdrive is that write-up, and to be clear up front, it's a concept. The repo holds a spec, a technical appendix, architecture and development docs, and a roadmap. There's no application code yet.

The idea

Classic defrag tools are reactive: they clean up after a disk gets messy. FlowDrive's pitch is to be predictive, learning which files you actually open and placing them before load times start to hurt.

The README aims it at gamers, video editors, developers and anyone juggling an SSD and an HDD by hand. If you've ever moved a library of folders to a new drive yourself, this Robocopy script is the manual version of what FlowDrive wants to automate.

The planned architecture

The design is four layers stacked on top of each other:

Layer Planned tech Job
User interface Python, PyQt6 Dashboard and controls
Intelligence layer Python, scikit-learn, TensorFlow or PyTorch, SQLite Learn access patterns, plan optimizations
Core engine C++ or Rust (undecided) Watch the file system, move and defragment files
OS layer Windows, Linux, macOS WinAPI and the defrag API on Windows, inotify and fanotify on Linux

The split is deliberate: disk work goes in a native core, learning and deciding go in Python where the ML libraries are, and the two talk over gRPC.

Note

The README's own status list says the spec and architecture are done, but the tech stack isn't finalized, the dev environment isn't set up, and the core engine and ML prototypes haven't started.

Design goals, not features

The README's Phase 1 list uses checkmarks, which reads like "done," but it's the MVP target, not shipped work: real-time access tracking, safe background defrag, SSD and HDD awareness, a simple dashboard and basic predictions. Later phases add better models, app-aware tiering and network storage.

The MVP bar puts safety first: zero data loss across 10,000+ file operations, under 5% CPU while it works, and at least 70% accuracy predicting file access.

Rough edges

This is a spec with a license on it, and the Quick Start says so: install instructions come later. Two docs the README links to, the API reference and the ML models page, aren't in the repo yet. The pricing ideas in there are thinking out loud, not a product you can buy.

Heads up

Nothing here touches your files today. A disk cleanup script will do more for a crowded drive right now.

What's next

First, pick C++ or Rust for the core engine; the README asks for input on exactly that. Then a prototype file monitor, since prediction means nothing without access data. If you want to see where the other experiments in this thread landed, DataChart is the one with the most code behind it.