GitGate is a git hosting platform built entirely on Cloudflare. No EC2 instances. No managed Kubernetes clusters. No RDS databases. Just Workers, D1, R2, Durable Objects, Queues, Workers AI, and Vectorize — running a pure TypeScript git engine at the edge.
It works. People use it. But getting here meant solving six problems that nearly killed the project. Each one seemed impossible within Cloudflare's constraints until we found the approach that made it work. Here's what broke, why it mattered, and how we fixed it.
1. D1 has a 10GB ceiling — and repositories grow fast
The problem
Cloudflare D1 is a serverless SQLite database. It's fast, it's cheap, and it has a hard size limit of 10GB per database. For a git hosting platform, this is a crisis. A single monorepo with years of history can easily exceed 10GB of metadata — commits, trees, refs, issues, pull requests, review comments, CI results.
Traditional git hosts solve this by running Postgres or MySQL on beefy dedicated servers with terabytes of storage. We don't have that option.
Why it matters
If you can't store repository metadata reliably, you don't have a git host. You have a demo. The 10GB limit isn't just a theoretical concern — it's the kind of wall you hit exactly when your product starts gaining traction, which is the worst possible time to discover an architectural dead end.
How we solved it
We shard globally. Every organization gets its own D1 database. A lightweight routing layer in the Worker resolves org-slug to a D1 database binding at request time. The router itself uses a small "directory" D1 database that maps organizations to their shard.
// Simplified shard resolution
const shard = await directory.prepare(
'SELECT db_id FROM org_shards WHERE slug = ?'
).bind(orgSlug).first();
const db = env[`D1_SHARD_${shard.db_id}`];
This gives each organization a full 10GB of metadata space — more than enough for even large enterprises. If an organization somehow approaches the limit, we can split it further (per-repository sharding), though we haven't needed to yet.
The tradeoff is complexity in the binding layer. Each shard is a separate D1 binding in wrangler.toml, and cross-organization queries (like global search) require fan-out across shards. But for a git host, most operations are scoped to a single organization anyway, so the fan-out cases are rare.
2. Workers can't do SSH — but git clients expect it
The problem
Every developer's muscle memory says git clone git@github.com:org/repo.git. That's SSH. Workers run on HTTP. There is no TCP socket API in Workers. You cannot accept SSH connections.
This sounds like a dealbreaker for a git host.
Why it matters
Git supports two primary transport protocols: SSH and HTTPS. If you can only support one, you need it to be seamless. Developers who can't use their preferred workflow will leave before they've pushed their first commit.
How we solved it
We went all-in on HTTPS and made it the first-class experience. Git's smart HTTP protocol is well-specified — it uses two endpoints (/info/refs and /git-upload-pack or /git-receive-pack) to negotiate and transfer packfiles.
We implemented the entire git smart HTTP protocol in TypeScript, running on Workers. The Worker parses the git protocol frames, resolves refs from D1, assembles objects from R2, and streams responses back to the client. Authentication uses personal access tokens over HTTPS basic auth, which git clients handle natively.
The key insight: most developers already have HTTPS configured as a fallback. By making HTTPS authentication frictionless (token-based, no SSH key management), we turned the limitation into an advantage. No more "add your SSH key to your account" onboarding friction. You generate a token, paste it once, and git remembers it.
3. Packfile generation eats CPU — and Workers have limits
The problem
When a client runs git clone or git fetch, the server needs to generate a packfile — a compressed bundle of all the objects the client needs. For a large repository, this involves reading hundreds of thousands of objects, computing deltas between similar objects, and compressing the result into a single stream.
This is compute-intensive work. Cloudflare Workers have a CPU time limit (typically 30 seconds on the paid plan, 10ms on free). A large git clone can easily exceed that.
Why it matters
If cloning breaks for large repos, your platform is a toy. Enterprise teams with years of history need reliable clones, and they need them to complete — not time out halfway through.
How we solved it
We split packfile generation into two phases. The negotiation phase runs in the Worker: it figures out which objects the client needs by comparing the client's refs with the server's refs. This is fast — it's just ref lookups in D1.
The assembly phase is offloaded. Once we know the object set, we use a pre-computed pack index stored in R2. For common operations (cloning the default branch, fetching recent commits), we maintain cached packfiles in R2 that can be served directly — no assembly required. The Worker just streams the pre-built pack to the client.
For edge cases where we can't serve a cached pack (unusual ref combinations, shallow clones with specific depths), we generate the packfile in a Durable Object with extended CPU time, streaming chunks to R2 as they're produced and then redirecting the client to the R2 object. The Durable Object's alarm API handles cleanup of these temporary packs after 24 hours.
4. Real-time collaboration without traditional WebSocket servers
The problem
A modern git host isn't just a place to push code. It's a collaboration platform. Pull request reviews, issue discussions, CI status updates — all of these benefit from real-time updates. When someone comments on your PR, you want to see it immediately, not after a page refresh.
Traditional platforms run dedicated WebSocket servers (or use services like Pusher/Ably). We don't have servers.
Why it matters
Real-time feedback is the difference between a platform that feels alive and one that feels like a static archive. Code review workflows especially suffer without it — reviewers and authors talking past each other because neither sees the other's comments until they manually reload.
How we solved it
Durable Objects are the answer. Each repository gets a CollabRoom Durable Object that manages WebSocket connections for everyone viewing that repository. When a Worker processes an event (new comment, CI status change, push), it sends a message to the CollabRoom, which broadcasts to all connected clients.
The architecture is elegant: Durable Objects are single-threaded and location-pinned, which means there's no distributed state problem. The CollabRoom for a given repository is one object in one location, and all WebSocket connections for that repository route to it. No pub/sub infrastructure, no message broker, no Redis.
We handle the geographic latency issue by keeping the CollabRoom in the region closest to the majority of a repository's contributors (determined by historical access patterns). For globally distributed teams, the extra 50-100ms of latency to a distant Durable Object is acceptable for real-time notifications.
5. Search at scale with Workers AI + Vectorize
The problem
Developers expect to search code. Not just filenames — the actual content of files across repositories. GitHub's code search is one of its most powerful features. Building anything comparable requires indexing millions of files and serving sub-second search results.
D1's FTS5 extension can handle text search within a single shard, but semantic search — finding code by meaning rather than exact text matches — requires something more.
Why it matters
A git host without search forces developers to clone repositories locally just to grep through them. That defeats the purpose of a web-based platform. And keyword search alone misses cases where developers describe what they want conceptually: "the function that validates email addresses" or "error handling for database timeouts."
How we solved it
We built a two-tier search system. Tier 1 is keyword search using D1's FTS5, sharded per organization. When code is pushed, a Queue consumer extracts text from each file and updates the FTS5 index. This handles exact and partial text matches.
Tier 2 is semantic search using Workers AI for embeddings and Vectorize for vector storage. The same Queue consumer that indexes text also generates embeddings for code files using a code-optimized model. These embeddings are stored in Vectorize, a vector database built into Cloudflare's platform.
When a user searches, we run both tiers in parallel. FTS5 handles queries that look like code (parseJSON, handleError), while Vectorize handles natural language queries ("function that converts CSV to JSON"). Results are merged and ranked by a combination of text relevance and semantic similarity.
// Parallel search execution
const [textResults, semanticResults] = await Promise.all([
searchFTS5(db, query, orgId),
searchVectorize(env.VECTORIZE, query, orgId)
]);
return mergeAndRank(textResults, semanticResults);
6. Auth without passwords — passkeys on Workers
The problem
Passwords are a liability for a code hosting platform. A compromised password means compromised source code. We wanted to support passkeys (WebAuthn/FIDO2) as a primary authentication method, not just a second factor. But the WebAuthn spec is complex, and implementing it correctly on an edge runtime with no native libraries is non-trivial.
Why it matters
Security is existential for a git host. If developers don't trust that their code is safe, nothing else matters. Passkeys eliminate entire classes of attacks: phishing, credential stuffing, password reuse. They're the future of authentication, and we wanted GitGate to be there from day one.
How we solved it
We implemented the full WebAuthn relying party specification in TypeScript on Workers. The Web Crypto API available in Workers handles the cryptographic operations: verifying attestation statements, validating assertion signatures, and checking challenge responses.
The registration flow stores credential public keys in D1. The authentication flow retrieves the credential, verifies the signature using crypto.subtle.verify, and validates the authenticator data (sign count, user presence, user verification flags). Session tokens are issued to KV with TTL-based expiry.
The hardest part wasn't the cryptography — it was handling the edge cases in the WebAuthn spec. Different authenticators (YubiKey, Touch ID, Windows Hello, Android biometrics) return slightly different attestation formats. We had to test against every major authenticator type and handle the variations.
The result: GitGate users can sign up with just a fingerprint or face scan. No password to remember, no password to leak. And it all runs on Workers with zero backend servers.
The takeaway
Every one of these problems had a moment where the obvious conclusion was "you can't do this on Cloudflare." D1 is too small. Workers can't do SSH. CPU limits are too tight. No WebSocket servers. No native search infrastructure. No crypto libraries for WebAuthn.
And every one of them had a solution that was ultimately better than the traditional approach. Sharding gave us isolation. HTTPS-only gave us simpler onboarding. Cached packfiles gave us faster clones. Durable Objects gave us zero-infrastructure real-time. Workers AI gave us semantic search without Elasticsearch. Passkeys gave us passwordless auth without a password database to protect.
The constraints weren't obstacles. They were design pressure that pushed us toward better architecture. GitGate is a stronger product because of the problems that "broke" it along the way.