Sabtu, 3 Oktober 2026


Enterprise AI costs are often discussed in terms of model pricing.

How much does one million tokens cost?

Which model is cheaper?

Should we use a smaller model for simple tasks?

Those questions matter.

But there is an even simpler question organisations should ask first.

Why are we paying to answer the same request again?

As AI adoption grows, the waste does not always come from expensive models. It often comes from repetition.

The same policy document is summarised again.

The same product description is generated again.

The same knowledge-base question is sent to the model by hundreds of users.

The same context is retrieved, packaged into a prompt and transmitted repeatedly.

Every request may look small.

At enterprise scale, they add up quickly.

This is where AI caching becomes interesting.

Instead of sending every request directly to a model, an AI platform can check whether an equivalent request has already been answered. If the result is still valid and the requester is authorised to see it, the system can return the cached response without making another model call.

That sounds simple.

Traditional applications have been caching data for decades.

AI makes the problem slightly different because two prompts do not have to be identical to mean the same thing.

“Summarise our annual leave policy.”

“What is the company leave policy?”

“How many annual leave days do employees receive?”

Depending on the use case, these may all be close enough that a semantic cache can recognise the similarity and reuse an existing result.

That can reduce token consumption, latency and repeated retrieval from downstream systems.

It can also make an AI Gateway more than just a routing layer.

The gateway can decide whether a request should go to a model, whether it can be answered from a trusted cache, whether a cheaper model is sufficient, or whether the response needs to be generated again because the underlying information has changed.

But caching introduces a trust problem.

Imagine an executive asks an AI system to summarise a confidential acquisition document. The response is cached.

Later, another employee asks a similar question.

A badly designed semantic cache might recognise the similarity and return information the second user was never authorised to access.

The token saving would be excellent.

The security outcome would not.

That is why AI caching cannot simply be based on prompt similarity.

Identity, data classification, tenant boundaries, permissions, source freshness and context all need to be part of the cache decision.

A cached response should inherit the security conditions under which it was created.

The system should know who generated it, which information sources were involved, who is allowed to retrieve it, when it expires and what happens when the underlying data changes.

This connects directly to a broader architectural principle discussed in Security Architecture Is Where Trust, Controls and Resilience Come Together: connectivity alone is not enough. Architecture has to define why access is permitted, which controls enforce that decision and what happens when those controls fail.

AI cost optimisation needs the same thinking.

A cache is not simply a cheaper place to get an answer.

It becomes another trusted component in the architecture.

Done properly, caching can reduce repeated model calls, improve response time and make enterprise AI considerably more efficient.

Done badly, it can become a fast way to replay stale or sensitive information to the wrong person.

The goal therefore should not be to cache everything.

It should be to avoid unnecessary AI requests without weakening the trust model around them.

Because the cheapest AI request is not necessarily the one sent to the cheapest model.

It is the one you never needed to send again.

Next
This is the most recent post.
Previous
Catatan Lama