Be updated, subscribe to the OpenKM news

S3 Storage for Document Management: How to Scale a Document Repository

Written by OpenKM on 25 september 2026

A CIO managing a growing document repository is inevitably faced with a key question: what happens when local storage approaches the platform’s limit?

More documents mean more versions, more indexed text, more previews, more backups, and greater business continuity requirements. For an infrastructure manager, the challenge is not simply to add capacity, but to decide where the content is stored—or how it is distributed—and how the architecture can evolve without disrupting document management for users.

OpenKM addresses this with a new capability planned for OpenKM 8.2.9: optional support for S3-compatible storage as an alternative to the traditional file system

In simple terms, this change makes it possible to separate two concepts more clearly: document management and the physical infrastructure where documents are stored.

What is S3 storage and what does it bring to document management?

S3 support makes it possible to use object storage as the backend for storing document content managed by OpenKM.

Until now, the common model was a datastore based on a filesystem. With the new architecture, an organization can continue using that model or consider S3-compatible buckets when it needs greater flexibility to scale, distribute, or reorganize its infrastructure.

The key is the abstraction of the storage layer. OpenKM continues to manage documents, permissions, metadata, versions, searches, and processes, while the underlying infrastructure can evolve independently.

For end users, document management logic remains within OpenKM. For the infrastructure team, there is a new option for deciding where the content is physically stored.

From traditional storage to a more flexible document architecture

A document repository contains more than just files. In addition to original documents and versions, a platform such as OpenKM manages extracted text for searching and indexing, representations used for previews, thumbnails, and other auxiliary resources.

How S3 storage works in a document management system

The planned architecture makes it possible to use different buckets or prefixes for different types of information.

This separation is particularly useful when an organization wants to apply different policies depending on the purpose of the content. An original document may have different retention and availability requirements from an automatically generated thumbnail or extracted text used for search.

There are also prefixes associated with tenants and shards, making it possible to organize content in architectures with multiple repositories, customers, or storage partitions.

The result is a document architecture that is less dependent on a single physical file system.

Cloud document storage: cloud, private, and hybrid scenarios

S3 is often associated directly with Amazon Web Services, but the concept is broader.

The planned architecture will use the standard AWS SDK and will allow custom endpoints to be configured, including Amazon S3 and object storage infrastructures compatible with its API.

This does not mean that every S3-compatible platform is automatically certified by OpenKM. The final list of supported platforms must be confirmed before making specific commercial claims.

From an architectural standpoint, however, this flexibility opens the door to different scenarios.

Hybrid architectures for document management

An organization can keep OpenKM on its current infrastructure, use OpenKM cloud services for document storage, deploy object storage on private infrastructure, or combine different components in a hybrid model.

The goal is not to impose a specific model, but to expand the available options.

How to manage large documents with S3 storage

Enterprise repositories do not always work with small files. Plans, CAD files, technical documentation, digitized case files, or large document packages can become very large.

For these cases, S3 support includes multipart transfers.

Instead of sending a large file in a single operation, it can be split into several parts. If a transfer is interrupted, the entire process does not necessarily have to start again from scratch.

The architecture also includes mechanisms to cancel incomplete uploads and prevent leftover fragments or unnecessary data from occupying storage space.

Deletion, restoration, and purge processes have also been considered as part of the workflow to reduce the risk of orphaned objects remaining after content is deleted from OpenKM.

How to migrate a document repository to S3 storage

This is probably one of the most important questions for an organization that already has hundreds of thousands or millions of stored documents.

Adopting S3 should not mean starting from scratch. The architecture includes a migration tool from filesystem storage to S3, making it possible to move existing content progressively.

The process can be run incrementally, repeated to add documents created during the migration, and supported by integrity checks to validate the transferred content.

This makes it possible to plan a phased transition: an initial migration can copy most of the repository. A final synchronization can then include changes made during the migration process. Only after that is the backend switched in a controlled manner.

This distinction is important: copying content to S3 and enabling S3 as the backend are two separate operations.

Migration must be part of a controlled project that includes backups, error checking, version validation, download and search testing, and integrity verification.

Security when migrating documents to S3 storage

Moving documents to S3 does not remove the need for security controls.

The credentials OpenKM uses to access storage must be protected. The design includes storing these credentials in protected configuration and excluding them from certain diagnostic and support packages, for example, to reduce the risk of accidental exposure.

Backend policies should also be designed according to the principle of least privilege.

S3 should not be confused with a complete backup or disaster recovery strategy. Storing documents in object storage does not replace the need to define backups, restoration procedures, retention policies, and regular recovery testing.

The infrastructure changes; operational responsibilities do not. The infrastructure may be different, but the same operational responsibilities remain.

When should you use OpenKM with S3 storage?

This model is particularly relevant when a repository is growing continuously, when there is a cloud or hybrid strategy, when an organization wants to reduce its dependence on the local storage of a single server, or when different types of content require different infrastructure policies.

It may also be relevant for organizations that want their document management platform to evolve relatively independently from the provider or technology used to store the objects.

That does not mean S3 is always the right option. Storage costs, number of operations, data transfer, geographic latency, retention policies, recovery requirements, and business continuity requirements must all be analyzed for each project.

The decision should not simply be “filesystem or S3,” but rather which architecture is best suited to managing the organization’s growth and requirements.

Benefits of using OpenKM with S3 storage

The main innovation of S3 support is not a change in the way users work with OpenKM.

It is a new infrastructure option for supporting the repository.

For IT teams, it means more options for managing capacity and growth. For solution architects, it helps support private cloud and hybrid scenarios. For document management, it makes it possible to change the underlying infrastructure while keeping permissions, search, versioning, automation, and document control logic within OpenKM.

Is your document repository growing and do you need a more flexible storage architecture? Contact the OpenKM team to explore how to integrate S3 storage into your environment and define a migration strategy based on your organization’s volume, infrastructure, and business continuity requirements.

 

Hubungi kami

Pertanyaan umum

JBA Solutions Sdn Bhd

OpenKM in 5 minutes!