Writer introduces new AI model and upgraded harness to contain token costs | TechCrunch
Across the AI industry, users are becoming more conscious of just how expensive their deployments can be —and feeling a new urgency to cut costs. But while open-source models offer significantly lower per-token costs, it can be difficult to find the right model for a given job.
On Thursday, Writer, which offers AI tools and agents for marketers, launched a new flagship model called Palmyra X6, aimed at solving that problem for its users. Built as a post-training variation on Z.ai’s open source model GLM-5.2, Writer says the new system should provide deployment-ready capabilities at a much lower price. The company estimates the new model, combined with changes to the companies harness infrastructure, will cut costs for its customers by as much as 50 percent for basic tasks.
Together with the new model, the company also released significant upgrades to its standard agentic harness. Both features will be available to Writer clients starting Thursday.
“I think the enterprise is absolutely sick of chasing the next benchmark,” CEO May Habib told TechCrunch. “They want flattening cost, and it seems like nobody can deliver that.”
The new approach puts particular emphasis on complex, multi-step tasks, executed faster and with fewer tokens. And Writer sees harness optimization as a crucial lever towards making that happen.
A recent paper from Writer researchers lends credence to this approach, testing small changes in harness efficiency across multiple different models. The research found that, in many cases, changes in the harness were a more reliable way to reduce costs than model choice, with costs falling an average of 40% across their testing.
“The harness is the one component whose efficiency multiplies across every model an organization runs—present and future,” the researchers wrote.
For Writer’s clients, the experience is still model-agnostic: Palmyra X6 will sit alongside other Writer models or outside models imported through Azure or Amazon Bedrock. But Habib also sees the push to cut costs as driving a broader distrust towards major AI labs, who have a financial incentive to drive up token use.
“The cost explosion here is just unprecedented for customers, and so is the degree to which CIOs are giving up on the labs,” Habib told TechCrunch, adding that the AI labs “don’t deeply understand right how to help an enterprise get benefit from AI.”
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.
