An MCP file upload sent as a base64 tool argument gets retyped by the model. The base64 inside a tool argument is model-generated output: the original bytes never move, and the model has to reproduce an opaque encoded string in its tool call. The file only arrives intact if that reproduction is perfect. Sometimes it isn’t, and when it isn’t, nothing reports an error.
I maintain redmine-mcp-server, an open-source MCP server with 51 core tools, usually deployed over HTTP on a different host from the agent calling it. In September 2026 one of the project’s most active contributors, @andilem, filed the cleanest reproduction of this I have seen: a 3,232-byte PNG that arrived as a broken image while every layer reported success. They also built the fix. I reviewed it.
TL;DR: MCP has no client-to-server file transfer, so a base64 upload is copied out by the model. A 3,232-byte PNG needed 4,312 characters and the tool call carried 4,280; the server stored a corrupt file and returned success. Move the bytes over a single-use HTTP upload ticket so the model only handles a path and a UUID, and verify a
sha256on whatever base64 remains.
If a file has to pass through a tool argument, you are not uploading it. You are asking a language model to type it.
The common failure pattern
Most HTTP-deployed MCP servers that accept files look like this:
file on disk -> base64 -> tool argument -> server -> API
It looks like a pipe, but the third step isn’t one, because there is no way to pipe a file into a tool argument. The agent runs base64, the output lands in its context, and then the model writes the tool call, reproducing the payload from what it just read.
For a 200-byte CSV, you might get away with it. For a real file it is a transcription task thousands of characters long, and nothing checks the result.
What actually arrived
The report is issue #305. An agent was asked to attach a small wireframe mockup to an issue, on server version 2.15.0. It did everything right:
1. base64 -w0 mockup.png > tmp correct
2. cat tmp 4312 chars, correct
3. update_redmine_issue
uploads=[{content_base64}] 4280 chars
In step 2 the transcript shows the full, correct payload in the tool output, head and tail intact. So the corruption happened between the model reading it and writing it back.
Thirty-two characters went missing. The first 1,209 characters matched the source and so did the last 2,459. The damage sat between offsets 1,209 and 1,334, with two isolated substitutions further on:
correct: ...MHQMHUmSoYO+hY+hIkgwdQwdDR5Jk6g6GDqSJEPH0DF0...
sent: ...MHQMHUmSoWPoYOhIkgwdQ8fQkSQZOhg6ho4kyDF0JEmG...
The source string explains why. A wireframe is mostly flat colour, and flat colour compresses into a near-repeating pattern. That stretch is where the model lost its place: insertions, deletions and transpositions, the same errors a person makes copying a long serial number by hand.
What landed in Redmine was 3,208 bytes against the source’s 3,232. It was identical up to byte 907, diverged inside the PNG’s IDAT stream, and had no IEND chunk. The payload was still valid base64 and under the size cap, so my server had nothing to object to. It stored the file, Redmine accepted it, the tool returned success, and the issue now carried a broken image.
Base64 upload trouble usually gets filed as a large-file problem. This happened at 3 KB, and a screenshot or a PDF is ten to a thousand times longer.
Why nothing inside MCP fixes it
The obvious move is to look for the protocol feature that handles this, and there isn’t one yet. Roots tell a server which directories exist on the client but grant no read access. Sampling and elicitation move text. A source_url pointing back at the caller fails twice: an SSRF guard rejects loopback and private ranges by design (see the SSRF fix in my MCP security audit), and across hosts a caller-side URL does not route anyway.
The long-running spec discussion on passing files from the client opens by calling base64 in tool calls “unreliable, particularly as file sizes increase”, and a July 2026 reply there still states that the protocol “has no interoperable client → server file-upload primitive”. A proposal exists, SEP-2631, and it is still an open draft.
The symptom is not unique to my server either. A Claude Code issue on the Drive connector reports a 12,012-byte spreadsheet arriving as 8,447 bytes with a normal success response. The cause there was never established and the issue was closed as not planned, but from the outside it looks the same: the call succeeds and the file is corrupt.
I call this the transcription boundary. Anything that crosses into a tool argument has been transcribed by the model. Short, meaningful values survive that: an ID, a path, a sentence. Long opaque strings do not reliably survive it, and base64 is exactly the kind of long, opaque string a model should not be asked to reproduce.
The fix: keep the bytes off the model’s output path
The known answer to large MCP uploads is a side channel. FutureSearch describes the presigned URL version well, framed around context window cost. I think that framing undersells it. Cost aside, the side channel removes the model from the byte-transfer path.
My server already moved bytes past the model in the other direction. Downloading an attachment returns a URL under /files/{file_id}, guarded by an unguessable ID with a TTL, not the content. @andilem’s proposal was the mirror image, reusing the same storage, expiry and cleanup. It shipped in 2.16.0 as three steps:
create_upload_ticket -> upload_url, ticket, expires_at
POST upload_url -> upload_id, size, sha256
update_redmine_issue -> uploads=[{upload_id: ...}]
The middle step is one shell command:
curl -sS -H "X-Upload-Ticket: $TICKET" \
--data-binary @mockup.png "$UPLOAD_URL"
The model writes a path into a command and reads a UUID back. Nothing long passes through it, so nothing long can be miscopied.
The ticket is what makes this safe to hand to a model at all. It is single-use, valid for 15 minutes by default, bounded by a byte cap, stored only as a hash, and it authorises exactly one upload and nothing else. A wrong ticket and an unknown ID return the same 404, so the endpoint cannot be used to probe which IDs exist. If the ticket leaks into a transcript, the worst case is one junk file that expires on its own.
A side benefit: a 5 MB attachment used to cost about 6.7 MB of base64 in the conversation even when it worked. Now it costs a path and a UUID.
Make the base64 route fail loudly
You cannot delete the base64 route. A browser-based connector has no shell, so it cannot issue the POST. For content the model itself generated, such as a short CSV or an SVG it drew, base64 is still the only source that works there.
So the second half of the fix is a detector. Every uploads item now accepts an optional sha256 and size_bytes, verified after the content is resolved and before anything reaches Redmine. Trimmed from the server:
if expected_sha256:
wanted = expected_sha256.strip().lower()
actual = hashlib.sha256(content_bytes).hexdigest()
if wanted != actual:
return {"error": (
f"sha256 mismatch: declared {wanted}, "
f"received {actual}. The content did "
"not arrive intact. Send the file with "
"create_upload_ticket instead."
)}
It works because of one asymmetry: reproducing a 64-character checksum is far less fragile than reproducing 4,312 characters of base64, and even a tiny change to the payload produces a different digest. The tool description tells the agent to take the checksum from the same shell invocation that encoded the file.
It is only a detector. It turns a silently broken attachment into an error that names the working route, and the agent still has to go and use that route. The issue’s author said so in an edit note: their first draft led with the checksum, and they rewrote it because moving the bytes is the fix and the checksum is the fallback.
There is also an optional cap, REDMINE_MCP_CONTENT_BASE64_MAX_BYTES. I left it unset by default. Nobody knows where the safe ceiling is, since it depends on the model and on how repetitive the payload is, and a low default would break deployments that work today.
The description has to point at the safe route
A safe transport is not enough if the tool description steers agents away from it. Ours listed file_path first and recommended it for “a file that is already there”, which is only true when the server runs on the caller’s machine, a condition no model can check. Issue #303 has the field report: an agent holding a local mockup tried file_path, got refused, and told the user it could not attach the file, although the error it had just read named content_base64.
The sources are now listed caller-first, upload_id before everything else, and content_base64 is described as what it is good for: small content the caller generated. Agents read descriptions as instructions. I hit the same thing when evaluating an MCP server with an LLM: the order and wording of a docstring is part of your API.
What review caught
The side channel fixed the corruption, but it is also a new upload endpoint, so it needed a security review. Reviewing PR #306, I found three holes:
filename=".."escaped the storage directory.os.path.basename("..")returns"..", so the temp file landed in the parent directory, where the cleanup job never looks.- The download route served staged uploads without a ticket. An upload ID travels in a URL and reaches proxy logs, so it cannot double as a download credential.
- Multipart defeated the size cap. Starlette’s
request.form()buffers the whole body before your code can check its length. The route now takes a raw body only.
All three passed the happy-path tests.
Checking your own server
If your MCP server accepts files over HTTP, these are the questions I would ask of it:
| Check | Why |
|---|---|
| Does any file path require base64 in an argument? | That file is retyped by the model |
| Can a caller send bytes over plain HTTP instead? | Takes the model off the byte path |
| Does the base64 route verify a checksum? | Otherwise corruption reports success |
The failure is hard to see because every component behaves correctly: the agent encodes properly, the payload is valid base64, the server enforces its limits, and the API returns 200. Only the file is wrong.
A tool argument is something a model wrote. Give files a route the model never touches, and make the route it does touch prove what it carried.