AI's unbelievable 437× memory shrinkage
In 2023, the word "Unbelievable" took over a million bytes of AI memory. DeepSeek's V4.1 Flash needs 2,670. What changed, and why it makes AI cheaper to run.
Read article
Insights / Notes from the work
Notes on AI, data and building useful products. What worked, what did not, and what the evidence actually supports.
In 2023, the word "Unbelievable" took over a million bytes of AI memory. DeepSeek's V4.1 Flash needs 2,670. What changed, and why it makes AI cheaper to run.
Read article
GPT-6 Astra can deliver better results with fewer tokens. What does its opaque reasoning mean for AI costs, auditability and trust in business applications?
Read article
/ Archive
128k tokens are 96k words in English for ChatGPT 3.5 and 4. The ratio is estimated to be 0.75 words per token. However, the answer is not...
Read article
/ Archive
Today, we release a massive dataset for non-commercial use, i.e. research or personal projects. The dataset covers Amazon product data for all of...
Read article
/ Archive
Tax Shrink is a new online tool that helps owner-operators of Limited companies in the UK calculate and visualise the ideal salary-to-dividend...
Read article
/ Archive
Large-language models (LLMs) are great generalists, but modifications are required for optimisation or specialist tasks. The easiest choice is...
Read article
/ Archive
Recently, OpenAI released GPT4 turbo preview with 128k at its DevDay. That addresses a serious limitation for Retrieval Augmented Generation (RAG)...
Read article
/ Archive
Today, I received access to the new custom GPT feature on ChatGPT, and it appears to do what Sam Altman demonstrated. The implications are...
Read article
/ Archive
OpenAI's DevDay announcement yesterday addresses issues I wrote about in the infeasibility of RAG after building Llamar.ai this summer. Did I get...
Read article
/ Archive
Over four months, I created a working retrieval-augmented generation (RAG) product prototype for a sizeable potential customer using a...
Read article