CrofAI promised the newest models at the lowest prices, often below the best offer on OpenRouter. The founder claimed to use custom inference engines and dismissed competitors as victims of "skill issues." The investigation published by Kendell.dev tells a different story: CrofAI was an OpenRouter wrapper that silently routed requests to cheaper or weaker models than the ones customers paid for.
One example from the report: expensive Kimi K3 was sold at $2 input and $10 output, but requests ended up on GLM 5.3 Flash via OpenRouter. In practice, that meant a 13.3x markup on input and 20x on output compared with the model actually served. CrofAI's "proprietary model family" was equally fake: greg-2-ultra routed to GLM 5.2, greg-1-mini to Qwen 3.5 9B, while greg-2-super, greg-1 and greg-1-super replaced answers with Kimi K2.7 Code. In private messages, the founder admitted that the story of models created by him was a lie.
Hardware inconsistencies make the fraud even clearer. CrofAI claimed to run Kimi K3 on RTX Pro 6000 cards rented through Vast. But Kimi K3 requires around 802 GiB even with Q2_K quantization, an aggressive level. The largest RTX PRO 6000 machine on Vast has eight cards, for a total of 765 GiB: not enough. In another statement, the founder said he was running deepseek-v4-flash-0731 on a local DGX Spark. A DGX Spark has 128 GB of memory, so it cannot host that model.
CrofAI's response to the report was first to announce the service shutdown and promise refunds. Then, around 4:30 AM UTC on September 15, the owner published a now-deleted post pretending that a "team" had taken over, claiming the founder's statements were made under stress. On X, each reply began with "Hey, Nathan here." The act lasted a few hours: after the public reminded him of the wire fraud allegations, the founder wiped his entire online presence. The domains nahcrof.com and crof.ai return 404, the X account is gone, and the subreddit is now private.
This story is more than a scam. It reveals a structural flaw in the low-cost cloud inference market: price per token has become the marketing unit, but it does not include verification of the model actually served. A provider can advertise one LLM and serve another, and the average user lacks simple tools to notice. The investigator documented five attempts by CrofAI to hide OpenRouter fingerprints, all failed: a check most companies do not perform.
Those chasing the cheapest tokens fund exactly this scheme. Legitimate providers, which bear real GPU and operating costs, are undercut by someone selling below cost because they are reselling someone else's service at a markup. When the price is too low compared with known hardware costs, one of two things is false: the price or the model.
Hardware inconsistencies are the acid test. Anyone evaluating on-premise or self-hosted deployment knows that VRAM and memory constraints are known and verifiable. This case shows that even a cloud provider can describe implausible architectures. For third-party API users, reducing risk means asking for evidence: model size, quantization, declared hardware, and an endpoint with verifiable logging. For those evaluating on-premise deployment, AI-RADAR offers analytical frameworks at /llm-onpremise to assess these trade-offs.
The question of refunds remains open. CrofAI promised to return money to those who asked, but with the domains returning 404 and the X account deleted there is no official channel left. It is the coherent ending of a provider that sold the fiction of low-cost inference until the hardware math caught up.
💬 Comments (0)
🔒 Log in or register to comment on articles.
No comments yet. Be the first to comment!