Why your data costs are rising

Regular AI use in organisations rose from 55% in 2023 to 88% in 2025. 

The reality: We aren't buying AI like software, we are consuming it like electricity. Every prompt, every response, every workflow uses tokens.

The cost is the sum of them all.

Long chats, large documents and unclear prompts use more tokens and can make answers less focused. Simple tricks can help you use tokens more effectively, get clearer answers and save time.  This is why better results do not always come from using more AI. They come from using it wisely.

LEARN HOW TO ACTIVELY REDUCE YOUR TOKENS

 

What is a token, in real terms

A token is a chunk of text, not a word, not a sentence: a unit the model uses to process and generate language.

Input tokens = what you send
Output tokens = what the model returns
Total tokens = what you pay for


Most providers price these separately, with output typically costing more than input. 

That one detail drives everything. 

AI Token Cost Calculator

Estimate monthly and annual AI spend across your workforce. Pick a model, set your headcount and expected prompts, and switch to the local hosting tab to compare.

Cloud LLM pricing
Token-based model. NZD pricing shown per 1M tokens.
250
Prompts per user per month
Credits per prompt (auto)

Credit bands from Microsoft Copilot Cowork task type guidance. For token-based LLMs the calculator translates each prompt class into representative input and output token volumes.

Estimated cost
Tokens / month
0
Cost / employee / month
NZD$0
Total cost / month
NZD$0
Annual projection
NZD$0
Cloud vs. local hosting
Cloud monthly
NZD$0
Annual
NZD$0
Local monthly
NZD$0
Annual
NZD$0
Adjust inputs to see break-even analysis.
FactorCloudLocal / on-device
PrivacyShared infrastructureData stays in-house
LatencyNetwork dependentLow and predictable
Model qualityFrontier-gradeOpen weights, capable
ScalabilityElastic on demandCapped by hardware
Setup complexityAPI key and goHardware, ops, tuning

Pricing estimates are indicative only. Token model rates converted from current OpenAI, Anthropic and Google published USD pricing at 1.65 NZD/USD. Microsoft Cowork rate uses Copilot Studio pay-as-you-go list at USD$0.01 per Copilot Credit. Verify current rates with your provider before committing.

Use case one

Internal Q&A assistant

For policy lookups, onboarding support, internal process questions, and team knowledge access across the business.

Best where AI is used often by many people, because small interactions scale quickly.

12 interactions / day 600 input tokens 300 output tokens
Token pattern
Low complexity, high frequency
Primary outcome
Faster knowledge access

Governance advice: Keep prompts short, use retrieval instead of pasting full documents, and cap response length so everyday queries do not inflate token use.

Reduced search time Consistent answers
Use case two

Document summarisation

For finance, legal, operations, and leadership teams reviewing reports, contracts, board packs, and long-form documentation.

Best where teams need rapid extraction of meaning from long, dense information.

8 interactions / day 3000 input tokens 600 output tokens
Token pattern
Heavy input, lower frequency
Primary outcome
Quicker review cycles

Governance advice: Chunk large files, summarise in stages, and submit only the relevant sections rather than full source material.

Faster decisions Lower manual effort
Use case three

Customer support augmentation

For service desks and customer teams drafting replies, summarising tickets, and speeding up frontline resolution workflows.

Best where throughput matters, because repeated small actions compound into meaningful monthly usage.

25 interactions / day 800 input tokens 400 output tokens
Token pattern
Operational, high volume
Primary outcome
Faster response handling

Governance advice: Standardise prompt templates, control output length, and reserve premium models for complex escalations only.

More consistent service Lower frontline pressure

The shift NZ leaders need to understand

Cloud wasn’t billed like this. Software wasn’t billed like this. AI is variable, behaviour-driven and difficult to forecast without modelling.

Azure already positions this under consumption-based cost management frameworks. Claude and other models reinforce the same pattern with token-based pricing tiers. Copilot adds another layer with metered usage inside productivity tools.

What's the real risk of consumption based cost management frameworks? 

AI usage expands quietly. Unmonitored, often with lack of training or education. Increased, long chat prompts. That combination creates three failure points leading to uncontrolled growth:

No visibility. No governance. No optimisation.

How to actively reduce token consumption across your organisation
 

01
Cut input tokens first

Most cost sits in what you send. Large prompts and duplicated context quietly inflate usage at scale. The rule? Be concise and specific.

  • Remove unnecessary context
  • Structure prompts instead of pasting raw content, ie using 'Context' 'Task' and 'Outcome' directives
  • Use retrieval (Projects, Notebooks, Documents etc) instead of full pasted copy injection
  • Implement Prompt Caching from specific prefixes
OUTCOME

Lower cost, faster responses, cleaner outputs.

02
Control output length

AI will generate as much as you allow. Without constraints, you pay for unnecessary verbosity every time. Add instructions to your LLM's memory.

  • Set explicit response limits
  • Ask for bullet points instead of essays
  • Avoid expand patterns unless required
OUTCOME

Reduced output tokens without compromising value.

03
Design for reuse

Organisations repeatedly pay to generate the same answers. Without reuse, token consumption compounds.

  • Cache common outputs
  • Store summaries once and reuse often
  • Build internal knowledge layers
OUTCOME

Shift cost company wide from generation to retrieval.

04
Segment use cases by model

Not every task needs a premium model. Using the wrong model for routine work drives unnecessary spend.

  • Use smaller models for repetitive tasks
  • Reserve advanced models for high-value outputs
  • Align model choice to business impact
OUTCOME

Immediate cost reduction with no user impact.

05
Introduce governance early

AI scales without friction. Without governance, usage grows invisibly until cost becomes embedded.

  • Set usage baselines per team
  • Track tokens per workflow
  • Align spend to measurable outcomes
OUTCOME

Visibility, control, predictable cost.

Change the behaviour, you change the cost. 

AI adoption does not fail because of technology. It fails because usage grows without visibility, control, or clear ownership across the business. Teams adopt tools, workflows evolve, and token consumption becomes embedded before governance catches up. This is not a model problem. It is an operating model problem.

Softsource works with New Zealand organisations to assess AI readiness, audit real usage patterns, and design governance models that bring cost, risk, and value back into alignment. 

01

AI readiness and rollout planning 

02

Token usage auditing and cost visibility 

03

Governance frameworks and control layers 

04

Model selection and workload optimisation 

05

Policy, security, and compliance alignment

06

Change management and adoption strategy 

Book a scoping session with the Softsource AI governance team to understand where your organisation is today, where cost is being created, and how to implement practical controls without slowing down adoption.

Back to Articles