Skip to content

Repository files navigation

protoc.protobuf-net.dev

The service behind protobuf-net.dev's non-.NET language targets: C++, Java, Kotlin, Objective-C, PHP, Python, Ruby, and Google's own C#.

protobuf-net.dev is a static site that runs protobuf-net.Reflection in the browser through WebAssembly, so C# and VB.NET are generated locally and no schema is uploaded. The languages here come from protoc, which is a native executable and cannot run client-side — no maintained WebAssembly build of the compiler exists, only of the runtime. Offering those targets at all means running the binary somewhere, and this is that somewhere.

Everything else about the site is unchanged: only these targets involve a server, and only when the user picks one.

How it works

Layer What it is
src/worker.ts A Cloudflare Worker: CORS, a rate limit, a size cap, and a hostname that explains itself when you visit it. Workers are V8 isolates and cannot run a binary, so it does none of the work itself.
container/ A Cloudflare Container — a linux/amd64 image holding protoc and a small Go HTTP server that shells out to it. This is the part that can execute a native compiler.

A Worker cannot exec, but it can route to a container, so the split is not a design choice so much as the shape of the platform. The container is addressed as a single named instance: the work is stateless, so any instance would do, and pinning the name means one warm container rather than several idling at once. Cloudflare puts it to sleep after a minute of inactivity, and starting it again takes a second or two.

The API

POST https://protoc.protobuf-net.dev/generate
content-type: application/json

{
  "language": "cpp",
  "files": { "my.proto": "syntax = \"proto3\"; message Person { string name = 1; }" },
  "entry": ["my.proto"]
}

files is a virtual filesystem: every entry is written to a scratch directory and put on protoc's include path, so files can import each other. entry names the ones to generate code for — everything else is available only as an import. Omitting entry generates for all of them, which is rarely what you want once imports are involved.

The response mirrors the shape the site already uses for its local codegen, so both paths render through the same code:

{
  "files": [{ "name": "my.pb.h", "text": "…" }],
  "errors": [{ "isError": true, "lineNumber": 7, "columnNumber": 3, "message": "Expected \";\"." }]
}

A schema that does not compile is a 200 with diagnostics and no files — it is an answer, not a failure. exception carries anything that stopped the request being answered at all, and comes with a 4xx or 5xx. Diagnostics use the file names you sent, not paths on the server.

GET / says only that the host supports protobuf-net.dev, and robots.txt disallows everything. Both are served by the Worker without waking the container. The endpoint is deliberately not documented there: this host exists for one site, and describing the API to a browser only helps whoever is not using it through that site. The description above is documentation for whoever maintains this, which is a different audience.

Imports, and matching the browser

protoc ships its own google/protobuf/**, but the site resolves imports from the corpus embedded in protobuf-net.Reflection, which reaches wider: google/api/**, google/type/**, google/rpc/**, google/longrunning/** and protobuf-net/** — bcl.proto and protogen.proto. Without those, a schema that compiles happily in the browser would fail here for want of an import, and the two halves of one site would disagree about what a valid schema is.

So the image carries that corpus too, taken from the protobuf-net repository at the tag pinned in container/Dockerfile, and the include path is ordered deliberately:

  1. the files in the request
  2. protoc's own include/
  3. the protobuf-net corpus

Where protoc and the corpus both have a file — the well-known types — protoc's copy wins. Its generators are built against a particular descriptor.proto, and feeding them a slightly different one is a confusing way to discover the difference. The corpus fills the gaps rather than replacing anything, and container/main_test.go asserts the ordering so a refactor cannot quietly invert it.

The corpus is a superset of what the browser embeds, because it is taken from the repository rather than unpicked from the assembly. That direction is harmless: an import may work here that does not work in the browser, but nothing that works in the browser fails here.

Limits

Request body 1 MiB, checked at the Worker and again in the container
Files per request 64
Generated output 8 MiB
protoc runtime 15 seconds
Rate limit 60 requests a minute per IP

The container runs as an unprivileged user with no outbound network access, writes only to a scratch directory that is deleted before the response is written, and rejects file names that are absolute, contain .., or are not .proto. protoc is a C++ parser being handed untrusted input by anyone who finds the hostname, so it is worth being unambiguous about what it can reach.

Deployment

Push to main. GitHub Actions vets and tests the Go server against a real protoc, typechecks the Worker, then runs wrangler deploy, which builds container/Dockerfile with the runner's Docker, pushes the image to Cloudflare's registry and deploys the Worker in front of it.

Two repository secrets are needed:

Secret Where it comes from
CLOUDFLARE_API_TOKEN Cloudflare dashboard → My Profile → API Tokens → Create Token → Edit Cloudflare Workers
CLOUDFLARE_ACCOUNT_ID Workers & Pages overview, right-hand column

The Edit Cloudflare Workers template already includes Account → Containers → Edit, which is what the image push needs, so the token needs no editing after it is created.

The hostname

wrangler.jsonc claims protoc.protobuf-net.dev as a custom domain, and Cloudflare creates the DNS record itself on first deploy — there is nothing to add by hand, but the protobuf-net.dev zone does have to be on the same Cloudflare account. The apex stays on GitHub Pages and is not touched: only this subdomain is a Worker.

The first deploy is also what creates the container application, so it is slower than the rest and worth watching. Container deploys roll out gradually; wrangler containers list shows where it is.

Cost

Workers Paid ($5/month) is the floor, and it includes 25 GiB-hours of container memory a month.

Three settings decide whether it can ever be more than that, and all three are deliberately at their smallest useful value:

instance_type lite 1/16 vCPU, 256 MiB — what is billed while awake
max_instances 1 the Worker pins one instance; this makes it a ceiling too
sleepAfter 1m in src/worker.ts, not the Wrangler config — it is a property of the container class

Memory bills on what is provisioned while the container is awake, so the worst case is not "lots of traffic" but "awake all month": about $1.75 on top of the plan, and that needs a request at least once a minute, day and night. An idle month costs nothing beyond the plan.

lite is a measured choice rather than a cautious one. Generating C++ from descriptor.proto — 61 KiB in, 2 MiB out, the largest schema the site offers as a sample — takes about a second on 1/16 vCPU and peaks at 46 MiB of the 256; three of those at once peak at 65 MiB. An ordinary schema is done in milliseconds. Four times the CPU would make the worst case a quarter-second and the common case no different, at four times the standing cost.

Working on it

docker build -t protocd container      # builds the image, downloads protoc and the corpus
docker run --rm -p 8080:8080 protocd   # the container's own API, no Worker in front
curl -s localhost:8080/ | jq

npm install
npm run dev                            # wrangler dev: the Worker, with the container behind it
npm run typecheck

wrangler dev needs Node 22+ and a running Docker daemon; it builds and runs the container locally, so the first start is slow and every start after that is not.

The Go tests skip themselves when protoc is not on disk, which is the usual case on a dev box:

cd container && go test ./...                      # string handling only
docker build -f - container <<'EOF'                # everything, against a real protoc
FROM golang:1.24-bookworm
COPY --from=protocd /opt/protoc /opt/protoc
COPY --from=protocd /opt/corpus /opt/corpus
WORKDIR /src
COPY . .
RUN go vet ./... && go test -v ./...
EOF

Updating protoc

container/Dockerfile pins the version and the release checksum; both change together:

gh api repos/protocolbuffers/protobuf/releases/latest --jq '.tag_name, (.assets[] | select(.name|test("linux-x86_64")) | .digest)'

Worth reading the release notes rather than bumping blind — protoc's generated output does change, and a new major version has been known to drop a language (--js_out left in v21, which is why JavaScript is not on the list here).

Updating the corpus

PROTOBUF_NET_TAG in container/Dockerfile should track the protobuf-net.Reflection version the site pins in src/ProtoGen.Wasm/ProtoGen.Wasm.csproj. They are separate repositories and will drift; the cost of drift is an import that resolves in one mode and not the other, so it is worth bumping this when the site's engine version moves.

Licence

Apache-2.0, matching protobuf-net.

About

Runs protoc for protobuf-net.dev's non-.NET language targets: a Cloudflare Worker in front of a container

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Sponsor this project

Packages

Contributors

Languages