Seventy demonstrations ran between December 2025 and July 2026. All of them sat on BrainChip Akida neuromorphic silicon and a heterogeneous compute ontology assembled from IBM Spectrum Symphony, IBM Spectrum LSF, and IBM Storage Scale, and all of them are listed publicly with dates, platforms, and video where video exists. The market has begun to sense that a single chip is not the unit of design. Writing in Semiconductor Digest, Hussein Osman of Lattice Semiconductor argues that a primary processor needs true companions rather than assistants, that designers should treat each chip as part of a team rather than a lead actor and an understudy, and that companion architecture will shortly stop being optional. The instinct is right and the frame around it is too small, since the argument stays inside edge artificial intelligence from the title onward, which is precisely the boundary the seventy demonstrations decline to observe.
The count is the least interesting thing about them.
A single demonstration is evidence about one occasion. Somebody chose the conditions, something ran, and the audience saw a result. Seventy demonstrations across defense, capital markets, public health, critical infrastructure, information integrity, and enterprise AI infrastructure are evidence of a different kind, because the conditions kept changing and one thing did not. Missile defense, pathogen surveillance, freight rail economics, wildfire discrimination, scriptural interpretation, sleep staging, and a race car at Talladega share no domain expertise, no data shape, no latency budget, and no customer. What the seven share is a substrate. The finding that survives across all seventy is not a result at all. The finding is what never had to be rebuilt.
What repetition establishes that one build cannot
Any competent team can produce an impressive demonstration in a chosen domain. The demonstration proves that a configuration existed in which the thing worked, and a buyer who has watched a few of them learns to discount accordingly, because the interesting question is never whether the demonstration ran. The interesting question is what the second demonstration cost.
Program economics answer that question in a way no single artifact can. When the same ten AKD1000 chips that spent weeks reading American freight rail were fitted with small WiFi radios and taught to sense human presence through the perturbation of ambient radio, the chips did not change, the orchestration did not change, and the shared storage layer did not change. Cerebra learned an empty room with no labels and reported movement across two floors, and the engineering that made the second capability possible was the engineering already paid for by the first. Weeks later the same architecture was steeped in three interpretive traditions and set to read the binding of Isaac, the Shema, and the Genesis creation account, mapping where Jewish, Christian, and Islamic readings converge and where they genuinely part. Rail traffic, human presence, and hermeneutics have nothing in common above the substrate. Below the substrate they are the same problem, namely sparse events arriving continuously and a fleet of small learners that must agree about what the events mean.
Marginal cost is the measurement that matters, and across seventy builds the marginal cost of the next domain fell to the cost of the domain knowledge itself. Nothing in the platform layer needed renegotiation. A demonstration program is worth reading as an amortization curve, and the curve here bends hard and early.
The chip does more than the category description allows
Coverage of neuromorphic computing confines the category to always-on edge sensing, and one 2026 forecast places edge deployment at more than half the market on the strength of drones, wearables, and IoT sensors, with another framing the entire opportunity around sub-milliwatt workloads that survive months on a battery. Edge sensing is real, valuable, and far too small a box.
Dense vision inference is the first thing that does not fit. SymRail assigned ten AKD1000 chips to ten live railfan cameras from Rochelle, Illinois to Folkston, Georgia and classified nine railcar types at 91 percent accuracy with four-bit weights and activations inside one megabyte of on-chip memory, producing freight visibility comparable to satellite data services priced above fifty thousand dollars a year. FireMesh classified every hot pixel from GOES-18, VIIRS, MODIS, and GLM as wildfire or confounder in roughly three and a half milliseconds against eight kilobytes of weights, and the hard part was never the speed. The hard part was discrimination, separating a wildfire from a refinery flare, an agricultural burn, and a metal roof throwing back the sun, which is the difference between an alert worth sending and an alert worth ignoring.
Language models are the second thing that does not fit. BrainChip’s TENNs-LLM-1b, a 1.24-billion-parameter state-space model, serves through the same vLLM plugin that serves the classifiers. A billion-parameter byte-level model was decomposed by function and run across more than twenty simulated chips wired as the layers of one network, so that temporal mixing ran on state-space parts and channel mixing ran on convolutional parts and the hardware topology became the model. In July, fifty thousand facts distilled out of Kimi-K3, a 2.78 trillion parameter teacher, into a three-million-parameter state-space student answered at 99.7 percent accuracy in 131 milliseconds across eleven Akida chips.
On-chip learning is the third. Peter van der Made’s 2007 tone-recognition proof-of-concept was reproduced on production AKD1000 silicon at 100 percent accuracy and roughly 1.9 milliseconds per inference, then extended across the fleet, with each new tone acquired on the chip while the previous nine were retained. A day later his complete functional-brain architecture from Higher Intelligence, thalamus through hippocampus, limbic system, and cortex, with structural plasticity, glial pruning, neuromodulation, and a predict-sense-update loop, ran across ten chips and taught itself a Tennessee rail junction overnight across more than eighteen thousand classifications with no labels and no retraining run.
An always-on sensor category cannot hold dense classification, generative serving, and unsupervised structural learning at once. One product cycle has been mistaken for a boundary condition, and an analyst who models the ceiling as the market has decided in advance that the category is small.
A fleet behaves as one mind
Single-chip inference is where the market description stops, and the most consequential property in the program appears only above that level.
SymSEAL put twelve AKD1000 chips in each of ten helmets, 120 inference units under a thirty-watt envelope per helmet, and fused them into one perception layer shared by ten operators, with Symphony orchestrating placement and IBM Storage Scale carrying distributed consensus between operator nodes. The architecture scales to a hundred helmets and a thousand chips without a redesign, because nothing in it depends on the number ten. The same fabric later drove a squad of ten simulated Unitree G1 humanoids on real gait physics, with 110 chips feeding a distributed hive that read the formation, took warning from a drone at range, and moved the squad from patrol to contact by consensus.
Consensus is the mechanism worth naming precisely. Inside each synthetic cortex, many columns judge how surprising the present moment is and vote, a thalamus relays the votes into a local verdict, and a council fuses the local verdicts across sites while weighting each by confidence and surfacing divergence, so the moment one brain notices what the other two miss is treated as attention rather than noise. Three such cortexes were brought online in Washington, Dallas, and Pittsburgh and reached k-of-N consensus across the continent on real Akida silicon, federated by Symphony multicluster and Storage Scale multicluster, with no node in charge and meaning crossing the wire through shared state. SymConstellation then drove the same council through a contested-spectrum run that jammed a link, removed a node, and partitioned the mesh, and the workload kept completing because it rerouted, migrated to a reachable peer, and re-reached consensus.
A system with no head cannot be decapitated, and a system whose program is the shape of its own graph has no fixed home for an adversary to strike. Fleets fused into one continuously learning brain over shared state are a system rather than a part, and the distinction is the whole distance between a component sale and an architecture.
The standard interface holds
Fragmented tooling is named in the reports as the headline barrier, and one report states plainly that absent standards block integration and force parallel codebases. Among many chips commonly perceived as edge devices, the gap is real, but BrainChip’s MetaTF Akida SDK has transitioned to the point where it remains accessible and mature along with other substrates that function as a peer in the heterogeneous compute ontology. Every emerging substrate arrives less turnkey than the incumbent, and graphics processors themselves were awkward to program before the surrounding software matured. The difficulty of new substrate tooling is often merely a matter of perception, though many today might still say CUDA is not exactly an easy skill to master. However, in an age of AI-assisted coding, these problems are more addressable by companies than they have been in the past, speeding adoption. A real gap and a disqualifying gap are separate claims, and coverage across the demonstration program here collapses them.
Akida now answers vLLM as a first-class backend. vLLM owns the API surface, the batching, the sampling, and the scheduling, and an out-of-tree plugin hands each forward pass to the Akida runtime, so the same /v1/completions and /classify endpoints an application already calls are answered by event-driven, int8, milliwatt-class silicon. A dashboard, a retrieval pipeline, or a semantic router that already speaks vLLM routes work to neuromorphic chips with no change of its own, which is the entire point. One hundred models hot-swapped across the fleet at an average of eleven milliseconds each, because model load is decoupled from the serving engine.
The July work pushed the same property across geography. Akida Cascade submitted every request to a single OpenAI-compatible endpoint and let a Symphony placement policy decide, while the work was in flight, where each request would execute, filling thirty-six IBM Cloud nodes in Dallas first, moving to ten on-premises nodes in Pittsburgh carrying real AKD1000 silicon, and running the remainder in Washington. Fifty-six nodes across three sites carried one workload simultaneously, and across 120,000 requests every answer was byte-identical to computing it locally, whether it ran on neuromorphic silicon in Pennsylvania or in simulation in Texas. The client posted to an endpoint and received a completion, with no software development kit, no awareness of the workload management domain, and no code change.
Even under batch scheduling, Akida works the same way along with other peers. Under LSF, ELIMs expose each chip as a resource the workload manager can perceive and place, reporting whether it is free, which model is currently mapped, and which tenant last used it, which turns a batch workload manager into an informed one that pre-warms models and keeps the duty cycle high. LSF already carries semiconductor design chains and large-scale model training across thousands of nodes, and NVIDIA runs it in its own design chain today. So, neuromorphic silicon and the dominant graphics processor now share a workload management substrate.
A portion of the program runs on the Akida simulation stack rather than on the chip, and the point cuts in favor of the toolchain. A simulator with fidelity sufficient to port straight to silicon is a maturity signal, not the reverse.
The data center is not reserved
The most consequential omission in the analyst frame is the assumption that neuromorphic computing stays at the edge while graphics processors keep the data center.
Horizontal scale settled that question early. A 46-node cloud fleet held 8,832 concurrent inference contexts at 99.4 percent of ideal scaling, with per-core throughput converging within five percent and per-context memory within three tenths of a percent across three independent clusters. Scaling that cleanly across cloud regions is what data-center workloads are required to do, and the orchestration substrate underneath is the one already carrying national-laboratory supercomputing and large-model training.
Research systems supply the second proof. Intel’s Hala Point, a 1.15-billion-neuron machine built from 1,152 Loihi 2 processors at Sandia, establishes that the architecture reaches data-center scale in neuron count today. What separates a national-laboratory research system from a commercial tier is orchestration and serving, and orchestration and serving are exactly what a heterogeneous compute ontology supplies. A fleet that answers the standard endpoint, places under the standard workload managers, and federates across regions is a data-center citizen.
Availability is where the coverage blurs most and costs a buyer most. Loihi reaches a narrow set of research partners and cannot be purchased. NorthPole remains a laboratory prototype despite frequent secondary claims otherwise. The AKD1000 has been purchasable as a chip or a PCIe board since January 2022, and the AKD1500 reached production shipments in mid-2026 on a 22nm process under 300 milliwatts with on-chip learning. Every demonstration described here runs on hardware a robotics team or a defense integrator can hold today. An analytical frame that flattens shipping silicon and unavailable research into one emerging bucket erases the distinction a buyer most needs.
Where the market description diverges
Three misreads recur across the published reports, and each one narrows the category in a way the hardware does not.
The first is the milliwatt-edge ceiling, answered above by dense classification, generative serving, and on-chip structural learning running on the same silicon.
The second is immature tooling read as undeployable, answered by a chip that sits behind vLLM, runs under Symphony and LSF, and hot-swaps a hundred models in about a second of aggregate load time.
The third should stop a careful reader, because at least one 2026 report frames composition itself as the strategy buyers are abandoning, arguing that enterprises are retreating from multi-chip, multi-vendor approaches toward single-vendor integrated platforms. The claim inverts the direction serious architecture travels. The confusion is understandable, since NVIDIA is consolidating vertically from silicon through networking, systems, and software toward a single-vendor integrated platform. Akida enters as a peer inside a larger composition rather than as a rival platform. What gets composed is kinds of processor rather than vendors of one kind, namely graphics, conventional, quantum, mainframe, and neuromorphic under one workload management domain, so a buyer assembling a heterogeneous compute ontology is not shopping for a second neuromorphic supplier and has no reason to wait for one. The correct architecture is the ontology that binds many kinds of processor together, and the single neuromorphic processor available for purchase today is enough to take its place in one.
Composition is the source of the reach on display. Under one orchestration domain the program has run quantum, neuromorphic, graphics, conventional processors, and mainframe as peer resource tiers, including a portfolio pipeline in which an IBM Heron R3 processor encodes features, Akida classifies the market regime, a language model writes the risk narrative, and z/OS settles the result. Neuromorphic cognition has been wrapped in CKKS homomorphic encryption seeded by a quantum random number generator, so the substrate computes on data it never decrypts. Composition unlocks capability that no single kind of processor and no single vendor delivers alone.
Internal inconsistency in the reports is the tell. For 2026 alone, published market-size estimates for the same market run from roughly 125 million dollars to more than 8 billion, with intermediate figures scattered between. A spread of more than sixty times for the same twelve months is not a disagreement about data. Each firm draws the boundary of the category in a different place, and the boundary is the thing under dispute.
Bad analysis is not a neutral intellectual failing, because valuation tracks the frame. A firm modeled as a single-chip edge-sensor vendor is valued as a single-chip edge-sensor vendor whatever the underlying trajectory, and the error travels from report to model to price with nothing in the loop performing primary verification.
Flexibility is the finding
Read the seventy builds as a substitution experiment and a pattern emerges that no individual demonstration announces. Every layer of the stack was swapped at least once while the layers above and below stayed put.
The chip moved hosts. The AKD1000 runs on Intel N100 nodes, on AMD EPYC, on commodity desktops, in the M.2 slot of an NVIDIA Jetson Orin Nano, and on a RUBIK Pi 3 built around the Qualcomm QCS6490. Bringing up the Qualcomm host took two weeks of PCIe forensics with a cross-built instrumented kernel and found two faults in the Qualcomm root complex, an L0s exit the host could not complete and a host-side event 16.6 seconds into boot that tore down an electrically clean link. Both faults were platform integration and neither was silicon. Seating a sub-watt inference engine in a slot that drones, robots, and industrial vision systems already carry means the deployed fleet inherits neuromorphic perception with no change to mounts, cooling, or benchmarks.
The model moved freely. Models are small files, hot-swappable per request, stored once on shared storage and routed by the workload manager, which is why a hundred of them live on every node at once and why a domain switch costs eleven milliseconds instead of a redeployment. Model swap across three domains was demonstrated on the OSINT panel at thirty-one milliseconds for the full rotation.
The orchestrator moved between paradigms. Symphony carries the service-oriented case, LSF carries the batch case, and the same fleet absorbs queued overnight work while serving online inference by day. Multicluster federation carries the geographic case, and a placement policy carries the live case.
The coordination substrate stayed a filesystem, which is the least obvious and most load-bearing choice in the program. IBM Storage Scale, or GPFS, serves as shared model store, three-tier KV cache in place of Redis, spike archive, lock-free shared nervous system for cortical columns posting votes, and the memory across which three federated sites reach consensus. A distributed filesystem has no single point of failure to nominate, no service to keep alive, and no protocol for a new participant to learn.
The interconnect became a design surface. An inexpensive Gowin FPGA on a Sipeed Tang Primer held a room’s world model in its own memory, with two neuromorphic chips writing into that memory over Ethernet with no operating system and no processor in the data path, one classifying WiFi channel state for presence and one classifying camera frames for drones, and the fabric fusing both modalities before handing a verdict to a tactical display. A second build put 1,024 neurons on that fabric beside an AKD1000 and modeled neuromodulation, homeostasis, short-term plasticity, dendritic integration, lateral inhibition, and stochastic firing as one loop, so the fabric decided when to learn while the chip decided what had been seen. Across thirty minutes of railcam footage the fabric spiked 1,024 times inside a single train and exactly zero times across the other 1,679 frames with nothing tuned. Both builds are the companion-silicon argument executed rather than described, since the Semiconductor Digest case rests on a field-programmable gate array holding a compact model at milliwatts and handling detection ahead of a primary processor. The difference is where the companion sits. Confined to a sensor it is a power-saving sentinel, which is the whole of what the coverage imagines. Placed between processors and addressed as shared memory it becomes an interconnect, and an interconnect is a datacenter component whose addressing, control plane, and routing carry straight to a fleet of neuromorphic hosts under the heterogeneous compute ontology.
The governance layer became data. Rules of engagement were expressed as Palantir Foundry ontology objects, queryable, auditable, version-controlled, and propagated to the edge, and an engagement proceeded only on N-of-M cryptographic confirmation from independent sensors, with one scenario ending in refusal because a single radar could not satisfy the provenance requirement. The same self-assembling ontology later drove Anduril’s Lattice, a production command and control platform, where live neuromorphic classifications from the Symphony cluster positioned entities on the operator’s map and every one of them carried provenance attribution back to Foundry. Bridging the two platforms took about three hours, which is one of the plainest measurements of the seams anywhere in the program.
Flexibility of that kind is not a property of any component. Flexibility is a property of the seams between components, and seams are the thing an architecture either designs or inherits. Every substitution above succeeded because the interface on each side was standard, narrow, and stable, namely a model file, a resource the workload manager can perceive, a shared address, an endpoint, or an ontology object. A program that swaps every layer and keeps running has demonstrated its interfaces, and interfaces are the only part of a stack with a decade-long life.
How a company should think about a heterogeneous compute ontology
An ontology, in the sense that matters operationally, is the workload management domain’s model of what exists in the estate and what each thing can do. Perception feeds the model, which is what an ELIM is, a small reporter that tells the workload manager something true about a resource, whether a chip is free, which model is mapped, which tenant last held it, or how hot a node is running. Placement follows perception. Policy governs placement. Capability grows when the ontology learns a new kind of thing, and existing workloads do not notice.
GPU-centric schedulers cannot express the idea. Run:ai and Kubernetes model a graphics processor and its fractions well, and neither has a vocabulary for a neuromorphic tier that learns on the device, a quantum processor with a queue and a calibration window, or a mainframe transaction that must settle. An estate managed by a control plane that cannot name a tier will never place work there, and the tier will be bought as a science project and stranded as one.
Four questions separate an ontology from a procurement list.
Can a new silicon type join the estate without re-architecting the applications above it or writing plugins to accommodate something new? The test is not whether the tier can be installed. The test is whether an application that has never heard of the tier can be served by it. Akida Cascade is that test run in public, since the client posted to an OpenAI-compatible endpoint and received byte-identical answers from three sites and two substrates while learning nothing about any of them.
Does the new tier answer the interface the organization already speaks? A processor tier that requires its own software development kit, its own serving path, and its own operational runbook is a parallel stack wearing the costume of an upgrade. Nothing downstream changed is the acceptance criterion, and any vendor answer longer than that sentence is a roadmap rather than a capability.
Can policy move work between sites and substrates without a human trigger? Elasticity that requires a person to notice, decide, and act is a runbook, not an architecture. Workload has moved under policy from Pennsylvania to Washington during a local surge and repatriated when pressure cleared, with no human in the path.
Is the coordination substrate a filesystem or a service? A coordination service is a component to operate, secure, scale, and mourn when it fails. Shared storage with real consistency semantics is already deployed, already backed up, and already understood by the operations team.
Sequencing follows from the questions. Start behind the interface the applications already call, so the first neuromorphic tier is invisible to everything above it. Add the tier to the ontology’s vocabulary before adding it to the floor. Move one workload class and benchmark cost per unit of work on both substrates, since classification, scoring, signal generation, and anomaly detection are what a grid already runs and where the efficiency gap is decisive. Measure the utilization of the estate rather than a mere benchmark of the new peer with tools designed for others. The strategy demonstrates the real return involved in adding a new peer such as Akida.
Economics reinforce the sequence. A task routed from a graphics processor to a milliwatt-class neuromorphic processor frees the graphics processor for work only it can do, so the same installed base absorbs more useful computation without an additional rack, an additional megawatt, or an additional cooling loop. Data-center electricity demand is projected to rise from 460 terawatt-hours in 2024 to more than 1,000 by 2030, individual campuses now draw between 500 megawatts and a full gigawatt, and interconnection queues hold years of pending requests. So, the binding constraint in the build-out has moved from silicon to power. Chips can be fabricated in twelve to eighteen months and grid capacity cannot. Every proposed remedy asks the grid, the regulator, or the launch vehicle to change first, while efficiency gives and asks nothing in return.
Governance belongs in the same design conversation and not in a later compliance review. Policy expressed as versioned objects, provenance carried cryptographically through the decision, encryption held at the inference boundary, deny-by-default posture under uncertainty, and refusal treated as a legitimate outcome are architectural properties. A continuous authentication build made the point plainly by composing an owner’s sleep signature on one chip, a real-versus-synthetic voice verdict on a second, and a rotating code as the one revocable factor, with the bias under uncertainty always set to refuse and a sleeping owner present but authorizing nothing.
The next front
An early demonstration put the return of an Akida-included heterogeneous compute ontology in plain numbers, and the arithmetic has only improved since February. A 500,000-core institutional trading grid dedicates roughly half its cores to classification, namely market-regime detection, signal generation, risk classification, and anomaly detection, every one of them high volume, low complexity, and tolerant of milliseconds. Fifteen thousand Intel N100 nodes carrying Akida M.2 cards absorb that entire load at about 100 million inferences per second. Boards for the neuromorphic tier run 279 dollars a node at current pricing, 150 dollars for the host and 129 for an AKD1500 M.2 module, so 4.19 million dollars of silicon and about 6 million dollars once chassis, networking, integration, and spares are counted. The 250,000 cores handed back are worth 125 million dollars at 500 dollars a core. A 6 million dollar investment therefore liberates more than 100 million dollars of compute the operator already owns, a return near twenty to one, and optimization and execution capacity doubles without the purchase of a single additional core.
Power moves the same direction. 500,000 cores at 10 watts draw 5 megawatts, and cooling at 40 percent overhead brings the facility to 7 megawatts, or 7.36 million dollars a year at 12 cents per kilowatt-hour. The redesigned grid draws 2.5 megawatts across the remaining cores, 150 kilowatts across the neuromorphic tier, and 1 megawatt of cooling, reaching 3.65 megawatts and 3.84 million dollars a year. Annual electricity falls by 3.52 million dollars while analytical throughput doubles.
Carry the same arithmetic into a data-center build and the shape holds at larger numbers. Take an inference estate of 250 million dollars in graphics-processor capital, and take 40 percent of the inference cycles to be work the neuromorphic processor runs for less. Routing that share onto a neuromorphic tier frees 100 million dollars of installed capacity for work only a graphics processor can do. Absorbing the load takes on the order of 15,000 chips, which prices between 5 and 15 million dollars depending on whether the chips ride commodity hosts or server-class ones with integration, spares, and support counted in, so the capital returns between seven and twenty times over. At 300,000 dollars for an eight-way graphics-processor node, the freed 100 million dollars is about 330 nodes, 3.3 megawatts of load, 4.3 megawatts once cooling is counted, and 4.55 million dollars a year of electricity. The neuromorphic tier that takes over the function draws 375 kilowatts and costs about 512,000 dollars a year to run, so roughly 4 million dollars of annual power comes back alongside the capital, more than 20 million dollars across five years, with no additional rack, megawatt, or cooling loop purchased.
An evaluation committee at a large institution will not accept list prices, and it should not. The ratio survives symmetric discounting, since an organization negotiating thirty percent off neuromorphic hardware negotiates thirty percent off graphics processors as well. Divergent discounts do move the ratio, and the larger order earns the deeper discount, so the honest version applies twenty percent to the small purchase and thirty-five percent to the large one. Loading the neuromorphic side fully, with factory integration, switching, cabling, rack space, buyer-side engineering, vendor support, and a full-time operator, gives a five-year net present value of cost near 10.1 million dollars at a nine percent discount rate. Support arrives the way support always arrives, since the buyer purchases a configured system from a server vendor and the contract covers the whole chassis at that vendor’s standard percentage of system price, with the neuromorphic modules inside the boundary exactly as the network card and the boot drive are. One contract, one vendor, one entry on the approved-supplier list, and no separate negotiation for the unfamiliar part. The other side of the ledger carries the graphics-processor capital never spent, the power and floor space never occupied, and the vendor support never contracted, which at twelve percent a year of a 300,000 dollar node runs to more than twice the power line. Net present value of the benefit lands near 108 million dollars, a return near eleven to one after every discount and every loaded cost, with payback on the tier inside six months. Priced the other way, as two costed programmes rather than as an avoidance, the conventional build that absorbs the same growth carries a five-year net present value near 176 million dollars against 12 million for the neuromorphic path, so the tier costs about seven percent of what it displaces, brings 94 kilowatts online instead of 3.3 megawatts, and occupies twelve racks instead of thirty-three. The comparison is like for like, since both cases absorb the same demand, which is inference and the work a production grid already runs, namely classification, scoring, signal generation, risk-regime detection, and anomaly detection, with the general compute around them left to capacity the institution already owns. The forty percent is an efficiency frontier rather than a census of eligible work, because sparse methods express the same workloads a grid already carries, so the question is never whether something can run on a neuromorphic processor but which substrate runs it for less, and the frontier moves outward as models, toolchains, and silicon improve.
The floor case deserves more weight than the headline. Freed capacity is worth something only when work is queued behind it, so an estate with no growth ahead of it books no avoided capital and no avoided support, and the return falls to the power and the space of the routed workload alone, roughly three to two across five years. Cost avoidance counts when the purchase already sat in the plan, and a committee is right to ask for the plan. Two sensitivities dominate the rest. Halving the crossover share from forty percent to twenty halves the case, and the density of the neuromorphic tier, meaning how many chips share one host, moves the answer further than the price of a chip does.
Two dollar lines come out of the same arithmetic and only one of them is a cost saving. Against a sustained hundred million inferences a second, the conventional build carries capacity at about 180,000 dollars per million inferences per second per year and the neuromorphic tier carries it at about 12,000, so identical work runs at fifteen times less per unit and every product consuming that capacity improves in margin by the difference. Hold the budget constant instead of the workload and the same ratio reads the other way. An institution already spending eighteen million a year to run ten thousand instruments against ten models apiece, continuously, buys with that same approved budget either a hundred and fifty models per instrument or a hundred and fifty thousand instruments at ten models each. Work worth more than 12,000 dollars a unit and less than 180,000 has moved from uneconomic to obvious, and the band between the two numbers is the frontier. Continuous monitoring replaces sampling, per-instrument models replace shared ones, and overnight batch becomes real time, not because anyone found a new algorithm but because the cost of running one crossed under the value of running one. The February case priced the revenue side directly, taking a conservative ten percent uplift on a five hundred million dollar book from doubling analytical throughput, which is fifty million a year and a present value near 194 million against a tier costing twelve. Sizing the category by what neuromorphic processors displace measures only the first line and misses the second, which is the larger of the two and the reason the exercise is properly called a frontier rather than a saving.
The last question a committee asks is what failure costs. Failure costs the tier and nothing beyond it, because the workload falls back onto graphics processors that were never decommissioned, the workload manager and the filesystem were already in production, and no application ever learned which substrate answered. An architecture whose downside is the write-off of a seven million dollar tier against an estate of a quarter billion is a different proposition from one that asks for the estate to be rebuilt around it.
Two properties keep the arithmetic from being a spreadsheet exercise. Routing has to happen automatically, which a placement policy over an ontology that can name the tier provides, and the applications have to remain unaware, which answering the standard endpoint provides. Both are demonstrated above. Absent either one the freed capacity stays theoretical, because a person deciding case by case which workload belongs on which substrate cannot keep pace with the arrival rate of the work.
The argument so far concedes the analysts’ own frame, in which inference dominates the workload and the efficiency dividend is measured in inference cycles. The concession may understate the case.
Inference and training do not exhaust what computation does. Van der Made’s architecture, demonstrated on a production-sized fleet, describes a third mode, a predict-sense-update loop in which a system perceives continuously, is surprised by the novel, and accumulates understanding with no retraining run. A system of that kind neither trains conventionally nor merely infers against a frozen model. Learning happens on the device, in the moment, at milliwatts. Two chips playing Pong on a court made of the program’s own running source code showed the operational consequence, since each chip updated its own weights every frame and climbed from chance to roughly 99 percent within a few hundred frames, which turns on-chip learning into a placeable, migratable, checkpointable cluster workload where the deployed model is the trained model, with no training cluster, no model registry, and no retrain-and-redeploy loop.
The industry currently imagines large world models living inside one enormous model in one data center. The same third mode supports the opposite architecture, a world model assembled from many small minds that each learn a local region forever and carry skill instead of knowledge, with the heavy language layer left optional and on the ground. Twelve chips modeling and driving one race car point the same direction, ten of them each having taught themselves one slice of the vehicle’s physics while two raced it, one carrying an experiential memory refined lap after lap and one learning only how hard to push.
Should that mode mature, neuromorphic silicon would not compete for a slice of the inference budget. Spiking hardware would instead open a category of work that no current silicon performs at any power level, and the efficiency case, already decisive, would become secondary to a capability case.
Whether the third mode matures is a question the next seventy builds will answer. The seventy already finished have retired a different question altogether.
Almost none of the substrate is new. Storage Scale has held production data since 1998 and Symphony has carried institutional grids for close to two decades, in clusters of hundreds of thousands of cores processing millions of tasks a second, inside at least twelve of the twenty largest global banks and across semiconductor design, risk, and research estates that have run continuously through several complete turnovers of the hardware beneath them. The integration between the two is older than most of the software an enterprise is currently paying to modernize. Seventy demonstrations added a tier underneath an interface that was already load bearing, which is a far narrower claim than a new platform, and a far cheaper one to accept.
Weigh a pilot against that record. A proof of concept exists to establish feasibility, and feasibility has now been established seventy times across eight months, twelve silicon architectures, six industry domains, and every substitution described above, which leaves demonstration seventy-one carrying almost no information. An organization that requires a pilot before committing is asking whether the substrate can carry the work, and the substrate has carried the work in public, with dates, platform lists, and video attached, on hardware available for purchase and under workload managers the organization very likely already owns. A seventy-first build inside a customer’s own domain would establish nothing the seventy do not, because the domain was never the variable under test. The substrate was the variable, and the substrate held every time.
Consider what the pilot actually costs, because the hardware is the smallest line in it. Inside a large organization a proof of concept consumes architecture review, security review, procurement, vendor onboarding, a data-sharing agreement, a segregated environment, and a year or more of calendar time, and it occupies the few people whose attention is the scarcest resource in the building. A failed implementation spends money that was budgeted for the purpose and leaves working knowledge behind. A proof of concept spends the one resource that cannot be re-budgeted, namely the attention of the people who would otherwise have built the thing, and it ends by producing information the public record already contained. Large organizations lose more to the ritual of proving what is already proven than they would lose to an implementation that failed outright.
A program of seventy builds of the heterogeneous compute ontology establishes something narrower and more useful than a forecast, namely that the substrate exists, ships, places work, serves the standard interfaces, federates across a continent, learns without supervision, and survives every substitution attempted against it. Across all seventy the differentiator was never the chip alone. The differentiator was the heterogeneous compute ontology around the chip, and the intersection where a fleet of neuromorphic processors behaves as one learning brain, speaks the enterprise’s standard interfaces, and runs alongside graphics, conventional, quantum, and mainframe tiers under a single workload management domain remains, at present, unoccupied.
The full list, with dates, platforms, and video, is at kevindjohnson.org/#demos, and the segmented analysis of the program by industry, by technology, and by market position is published alongside it.