SCRIPT LIBRARY · POWERSHELL
Count Manuscript Words by Chapter, Word Files Included
Get a word count for every chapter file and the whole manuscript, from Markdown, plain text or .docx, with a running total and a progress bar toward your target.
- What it does
- Counts the words in every .md, .txt and .docx chapter file in a folder, in natural chapter order, and returns one row per chapter with a running total, then a total row. Give it a target and you get the percentage and a progress bar too.
- Requires
- Windows PowerShell 5.1 or PowerShell 7+
- No modules, and Word doesn't need to be installed
- Permissions
- Read access to the chapter files. Nothing is changed.
- Runs on
- Windows 10/11, macOS or Linux with PowerShell 7
- Tested
- Parse-checked and run in PowerShell 7.4 against Markdown with front matter, comments, links and a scene break, a plain-text file, a .docx with a tracked deletion, and a corrupt .docx, with and without -Target
Part 2 of the thread Tools for the writing desk
If you write one file per chapter, "how long is the book?" means opening every file, or pasting everything into one document just to look at the counter in the corner. And you want to know more than the total: which chapter is running long, which is thin, and how far you are from the number you promised yourself.
My first version of this was a quick loop that split on whitespace and printed a list. It was fine for Markdown and hopeless for Word files, which it read as raw bytes and counted as gibberish. This one reads .docx properly (it's a ZIP of XML inside) and cleans Markdown up before counting, so formatting marks, links and your notes-to-self in comments don't pad the number.
Output is objects, one per chapter, so you can sort them, chart them, or drop them into a spreadsheet to track progress week by week.
<#
.SYNOPSIS
Counts words per chapter file and for the whole manuscript, with an optional target.
.DESCRIPTION
Reads every .md, .txt and .docx file in a folder, in natural order (Chapter 2 before Chapter 10), and returns
one object per file with its word count and a running total, then a Total row.
Markdown is cleaned up before counting: YAML front matter, HTML comments (handy for notes to yourself),
link targets, image tags and formatting marks are ignored, so **bold** and [links](url) count as the words
you'd read. Word documents are read straight from the .docx package (it's a ZIP of XML), so Word doesn't
need to be installed; tracked deletions, comments and footnotes aren't counted. A "word" is any run of
non-space characters with at least one letter or digit in it, so a lone dash or *** scene break doesn't count.
Give it a -Target and each row also shows the percentage reached, and you get a progress bar at the end.
.PARAMETER Path
A folder of chapter files, or specific files. Defaults to the current folder.
.PARAMETER Include
Extensions to count. Default: .md, .txt, .docx.
.PARAMETER Recurse
Include subfolders (for manuscripts split into part folders).
.PARAMETER Target
Your target word count for the whole manuscript.
.PARAMETER ExcludeTotal
Leave the Total row off, if you're feeding the output somewhere that sums it itself.
.EXAMPLE
.\Measure-ManuscriptWords.ps1 -Path .\chapters -Target 80000
.EXAMPLE
.\Measure-ManuscriptWords.ps1 -Path .\chapters | Export-Csv wordcount.csv -NoTypeInformation
#>
[CmdletBinding()]
param(
[Parameter(ValueFromPipeline, ValueFromPipelineByPropertyName)]
[Alias('FullName')]
[string[]]$Path = '.',
[string[]]$Include = @('.md', '.txt', '.docx'),
[switch]$Recurse,
[ValidateRange(1, 10000000)][int]$Target,
[switch]$ExcludeTotal
)
begin {
Add-Type -AssemblyName System.IO.Compression, System.IO.Compression.FileSystem
$Include = @($Include | ForEach-Object { if ($_ -like '.*') { $_.ToLower() } else { ".$_".ToLower() } })
$files = [System.Collections.Generic.List[IO.FileInfo]]::new()
function Get-DocxText([string]$File) {
$zip = [IO.Compression.ZipFile]::OpenRead($File)
try {
$entry = $zip.GetEntry('word/document.xml')
if (-not $entry) { throw 'No word/document.xml inside. Is this really a .docx?' }
$reader = [IO.StreamReader]::new($entry.Open())
try { [xml]$xml = $reader.ReadToEnd() } finally { $reader.Dispose() }
}
finally { $zip.Dispose() }
$ns = [Xml.XmlNamespaceManager]::new($xml.NameTable)
$ns.AddNamespace('w', 'http://schemas.openxmlformats.org/wordprocessingml/2006/main')
# One line per paragraph; w:t holds the visible text, w:tab and w:br separate words.
$paragraphs = foreach ($p in $xml.SelectNodes('//w:body//w:p', $ns)) {
($p.SelectNodes('.//w:t | .//w:tab | .//w:br', $ns) | ForEach-Object { if ($_.LocalName -eq 't') { $_.InnerText } else { ' ' } }) -join ''
}
$paragraphs -join "`n"
}
function Get-MarkdownText([string]$Text) {
$Text = $Text -replace '\A', ''
$Text = $Text -replace '(?s)\A---\r?\n.*?\r?\n(---|\.\.\.)\r?\n', '' # front matter
$Text = $Text -replace '(?s)<!--.*?-->', '' # comments and notes to self
$Text = $Text -replace '!\[[^\]]*\]\([^)]*\)', '' # images
$Text = $Text -replace '\[([^\]]*)\]\([^)]*\)', '$1' # links: keep the text
$Text = $Text -replace '<[^>]+>', ' ' # stray HTML tags
$Text -replace '(?m)^\s{0,3}(#{1,6}|>|[-*+]|\d+\.)\s+', '' -replace '[*_~`]+', ''
}
function Measure-Word([string]$Text) {
if (-not $Text) { return 0 }
@($Text -split '\s+' | Where-Object { $_ -match '[\p{L}\p{N}]' }).Count
}
}
process {
foreach ($item in $Path) {
if (Test-Path -LiteralPath $item -PathType Container) {
Get-ChildItem -LiteralPath $item -File -Recurse:$Recurse | Where-Object { $Include -contains $_.Extension.ToLower() -and $_.Name -notlike '~$*' } | ForEach-Object { $files.Add($_) }
}
elseif (Test-Path -LiteralPath $item -PathType Leaf) { $files.Add((Get-Item -LiteralPath $item)) }
else { Write-Warning "Not found: $item" }
}
}
end {
if (-not $files.Count) { Write-Warning 'No chapter files found.'; return }
# Natural sort: pad every number so "Chapter 2" sorts before "Chapter 10".
$sorted = $files | Sort-Object -Unique FullName | Sort-Object { $_.DirectoryName }, { [regex]::Replace($_.Name, '\d+', { $args[0].Value.PadLeft(10, '0') }) }
$running = 0; $i = 0; $counted = 0
foreach ($file in $sorted) {
$i++
Write-Progress -Activity 'Counting words' -Status $file.Name -PercentComplete (100 * $i / @($sorted).Count)
$row = [ordered]@{ Chapter = $file.BaseName; File = $file.Name; Words = $null; RunningTotal = $null }
try {
$text = switch ($file.Extension.ToLower()) {
'.docx' { Get-DocxText $file.FullName }
'.md' { Get-MarkdownText (Get-Content -LiteralPath $file.FullName -Raw -Encoding UTF8 -ErrorAction Stop) }
default { Get-Content -LiteralPath $file.FullName -Raw -Encoding UTF8 -ErrorAction Stop }
}
$row.Words = Measure-Word $text
$running += $row.Words
$counted++
}
catch {
Write-Warning "$($file.Name): $($_.Exception.Message)"
}
$row.RunningTotal = $running
if ($Target) { $row.PercentOfTarget = [math]::Round(100 * $running / $Target, 1) }
[pscustomobject]$row
}
Write-Progress -Activity 'Counting words' -Completed
if (-not $ExcludeTotal) {
$total = [ordered]@{ Chapter = 'Total'; File = "$counted file(s)"; Words = $running; RunningTotal = $running }
if ($Target) { $total.PercentOfTarget = [math]::Round(100 * $running / $Target, 1) }
[pscustomobject]$total
}
if ($Target) {
$pct = [math]::Min(1.0, $running / $Target)
$filled = [int][math]::Round(30 * $pct)
$left = [math]::Max(0, $Target - $running)
$bar = '[' + ('#' * $filled) + ('-' * (30 - $filled)) + ']'
Write-Host ('{0} {1:N0} of {2:N0} words ({3:N0}%){4}' -f $bar, $running, $Target, (100 * $running / $Target), $(if ($left) { ", $($left.ToString('N0')) to go" } else { '. Done!' }))
}
}
Parameters
| Parameter | Type | Default | What it's for |
|---|---|---|---|
-Path | string[] | . | A folder of chapter files, or specific files. Takes pipeline input. |
-Include | string[] | .md, .txt, .docx | Which extensions to count. |
-Recurse | switch | — | Include subfolders, for books split into part folders. |
-Target | int | — | Target word count for the whole manuscript. Adds a PercentOfTarget column and a progress bar. |
-ExcludeTotal | switch | — | Leave off the Total row, if whatever you're feeding it adds its own. |
Run it
Chapter counts and a bar toward 80,000 words.
.\Measure-ManuscriptWords.ps1 -Path .\chapters -Target 80000Longest chapters first.
.\Measure-ManuscriptWords.ps1 -Path .\chapters -ExcludeTotal | Sort-Object Words -Descending | Select-Object -First 5Append today's total to a progress log.
.\Measure-ManuscriptWords.ps1 -Path .\chapters | Where-Object Chapter -eq Total | Select-Object @{n='Date';e={Get-Date -Format yyyy-MM-dd}}, Words | Export-Csv progress.csv -Append -NoTypeInformationJust the Word files in a drafts folder.
.\Measure-ManuscriptWords.ps1 -Path .\drafts -Include .docxWhat you'll see
Chapter File Words RunningTotal PercentOfTarget
------- ---- ----- ------------ ---------------
Chapter 01 Chapter 01.md 3412 3412 4.3
Chapter 02 Chapter 02.md 2988 6400 8.0
Chapter 03 Chapter 03.docx 4105 10505 13.1
...
Chapter 18 Chapter 18.md 3890 61220 76.5
Total 18 file(s) 61220 61220 76.5
[#######################-------] 61,220 of 80,000 words (77%), 18,780 to go
How it works
- Collect the files. It gathers every file with an included extension from the folders you give it, skipping Word's
~$lock files, then sorts them with every number padded out, so the running total follows your chapter order. - Get plain text. Markdown has its front matter, comments, images, link targets, heading marks and emphasis characters stripped. A
.docxis opened as a ZIP, and the text is pulled fromword/document.xml: thew:truns, with tabs and line breaks treated as spaces. Deleted text lives inw:delText, so it's skipped automatically. - Count words. The text is split on whitespace, and only pieces with at least one letter or digit count. That keeps
* * *scene breaks and em dashes out of the total. - Return rows. Each chapter comes back with its count and the running total, plus
PercentOfTargetif you set-Target. A file that can't be read gets a warning and a blank count, and the rest carry on. - Draw the bar. With a target, the last thing it prints is a 30-character bar with the words left to go.
Take it further
- Build the book next. When the counts look right, Convert-MarkdownToKdpHtml turns the same folder of chapters into a single Kindle-ready file.
- Chart the pace. Log the Total row daily with the third example, then open the CSV in a spreadsheet for a words-per-day line.
- Catch lopsided chapters. Pipe the rows to
Measure-Object Words -Average -Maximum -Minimumto see which chapters are way off the average.
Things that'll trip you up
- It won't match Word's number exactly. Every tool counts a little differently. A hyphenated word like "well-known" counts once here, as it does in Word, but a dash standing on its own between spaces doesn't count at all, and tools disagree on that one. Expect to be within a percent or so of Word or Scrivener, which is plenty for tracking progress.
- Comments in Markdown don't count. Anything inside <!-- --> is ignored, along with YAML front matter. That's deliberate, so notes to yourself don't inflate the count. If you keep scene notes in plain text instead, they'll be counted.
- Tracked changes count as accepted. In a .docx, inserted text counts and deleted text doesn't, as if you'd accepted everything. Word comments and footnotes aren't counted at all.
- Name chapters so they sort. The natural sort handles "Chapter 2" before "Chapter 10", but a prologue named "Prologue" sorts after "Chapter". Prefix it with 00 if the running total should start there.