Wes Ellis./ a personal notebook
Technology. Stories. Side projects.
A few things worth writing down.
← Back to Engineering

Engineering

Chasing ClientOtherError in Azure Files: Is It Your SDK or Your Code?

THE SHORT VERSION4 points
  • ClientOtherError is the catch-all for ordinary 4xx responses (mostly 404, 409 and 412), and a lot of it is normal.
  • Split Transactions by API name, then query StorageFileLogs and read the UserAgentHeader to see who's calling.
  • In the code, look for legacy SDKs, an exists-check before every call, retries on 404s, and badly built paths.
  • A steady baseline is fine. A sudden spike, or one call on one path thousands of times a minute, is a bug.

Someone opens the metrics blade for a storage account, splits Transactions by response type, and sees a big pile of ClientOtherError. Cue the panic.

Before anyone rewrites anything: a lot of ClientOtherError is completely normal. The trick is figuring out which calls are producing it and whether your application (or the SDK it uses) is doing something it shouldn't.

What ClientOtherError actually means

Azure Storage sorts failed requests into buckets. The ones with their own names are things like throttling, timeouts, authorization errors, and network errors. ClientOtherError is the catch-all for the rest of the 4xx family, which mostly means:

Status Usually because
404 Not Found The file or directory isn't there.
409 Conflict The thing you tried to create already exists, or a lease is in the way.
412 Precondition Failed A conditional request's ETag check doesn't match.

Plenty of those are code working as designed. "Check if it exists, create it if not" generates a 404 on every first run. A CreateIfNotExists call on a share that's already there sends a create, gets a 409, and quietly reports success to your code. Storage counts that 409 as a ClientOtherError every time.

SMB clients are chatty in the same way. Windows Explorer and plenty of applications probe for files that aren't there (desktop.ini, thumbnails, temp files), and every one of those misses shows up in the metric.

So the question isn't "why do we have ClientOtherError?" It's "which operations produce it, and is that expected?"

Step 1: find out which calls are failing

Metrics give you the shape. Split the Transactions metric on the file service by API name and Response type, and you'll see whether it's GetFileProperties or CreateDirectory doing most of the damage.

For the details, turn on a diagnostic setting for the file service that sends StorageRead, StorageWrite and StorageDelete logs to a Log Analytics workspace. Give it some traffic, then:

StorageFileLogs
| where TimeGenerated > ago(24h)
| where StatusCode between (400 .. 499)
| summarize Requests = count() by OperationName, StatusCode, StatusText, UserAgentHeader
| order by Requests desc

Tip

The UserAgentHeader column is the one to look at. It tells you whether the noise is from SMB, a specific SDK and version, AzCopy, or something custom calling the REST API directly. Uri and CallerIpAddress narrow it down to a path or a machine.

Step 2: look at the code making those calls

Once you know which application and which operation, open the code that talks to storage.

Which library is it using?

Current SDKs:

Language Package
.NET Azure.Storage.Files.Shares
Python azure-storage-file-share
Java azure-storage-file-share
JavaScript @azure/storage-file-share

Heads up

If you find WindowsAzure.Storage or Microsoft.Azure.Storage.File, those are the legacy libraries and they're past end of support. Plan the upgrade whether or not it fixes this.

Is it probing before acting?

An Exists() before every read or a CreateIfNotExists() before every write doubles your transactions and fills the metric with 404s and 409s. It's often cheaper to just do the operation and handle the specific failure:

try
{
    await directoryClient.CreateAsync();
}
catch (RequestFailedException ex) when (ex.ErrorCode == ShareErrorCode.ResourceAlreadyExists)
{
    // Already there. Carry on.
}

That still produces a 409 when the directory exists, but only when it matters, not on every call.

Is it retrying things that will never succeed?

A retry policy that treats a 404 as transient will hammer the same missing file over and over. The SDK's built-in retry logic is sensible; custom wrappers around it often aren't.

Are paths built correctly?

Double slashes, stray URL encoding, or a missing share name will give you a steady stream of 404s. The Uri column in the logs makes these obvious.

Are leases and ETags handled?

412s and lease-related 409s usually mean two writers are fighting over the same file, or the code is holding on to a stale ETag.

When to worry

If ClientOtherError tracks your normal traffic and the operations make sense, you're fine. Leave it alone and maybe exclude it from alerting.

If it spikes suddenly, lines up with user complaints, or it's one operation on one path repeating thousands of times a minute, you've found a bug. The logs will tell you which app and which call, which beats reading an entire codebase looking for the problem.