How AI is applied across API Evangelist and APIs.io. Read my AI disclosure →
API Evangelist API Evangelist
Discovery
Learnings
Guidance
Toolbox
Alignment
API Evangelist LLC

WebHDFS

Apache Hadoop’s HTTP interface to HDFS, introduced in 2011 — a complete file-system vocabulary over REST (list, status, open with offset and length, mkdirs, rename, delete, permissions, owners, content summary) with a two-step write in which the namenode answers a CREATE or APPEND with a 307 redirect to the datanode that will take the bytes. It is the clearest production precedent for “send metadata first, receive an upload target, then send bytes,” and the documentation says exactly why it was built that way.

WebHDFS is what a POSIX-ish file system looks like when you put it behind HTTP and keep all the semantics. Every operation is a query parameter — ?op=LISTSTATUS, ?op=GETFILESTATUS, ?op=OPEN&offset=&length=, ?op=MKDIRS, ?op=RENAME, ?op=DELETE, ?op=SETPERMISSION, ?op=SETOWNER, ?op=GETCONTENTSUMMARY — and the responses are JSON. It was built so that clients in any language, and clusters that could not share a Java classpath, could reach HDFS.

  • The full vocabulary - Status, listing, ranged reads, directories, rename, delete, permissions, ownership, quotas and summaries. Nothing a shell user expects is missing.
  • Two-step writes - CREATE and APPEND are PUT to the namenode, which returns 307 Temporary Redirect with a Location on the datanode; the client then PUTs the bytes there. The docs explain that Expect: 100-continue would have done this in one request, but common HTTP libraries got it wrong, so the redirect is the workaround.
  • Gateways hold the credentials - Through Apache Knox or HttpFS, WebHDFS runs behind a gateway that authenticates the caller and talks to the cluster on its behalf, so the client never holds the cluster credential.
  • Copied by the clouds - Azure Data Lake Storage Gen2’s REST “Path” operations are the same shape with a lease and a flush step.

WebHDFS is in this catalog because it is the precedent nobody cites. When the Model Context Protocol Files Working Group agreed in October 2026 on “metadata first, upload target back, bytes, then finalize,” and separately required that uploads work behind a gateway that holds the OAuth credential, they described WebHDFS’s 307 and Knox. The Hadoop project solved both in 2011, wrote down the reason, and shipped it in every release since. It is also a good reminder that a file system over HTTP does not need new verbs the way WebDAV assumed; it needs a complete list of operations and an honest answer to where the bytes go.