Infinity startup Inference raises $15M from Touring Capital, OpenAI, and anthropology researchers


Amnesty International Infinity Infrastructure Company It announced a $15 million raise at a $100 million valuation on Monday from investors including Touring Capital, Principal VC and researchers from companies including OpenAI and Anthropic.

The startup is building software to make it easier for AI chips to run AI models. One of the main reasons Nvidia became the number one player is not only its high-performance chips, but also its CUDA (Compute Unified Device Architecture) software, which allows GPUs (originally designed to run graphics) to act as general-purpose CPUs. The largest AI development frameworks PyTorch and TensorFlow are built on top of CUDA. This allows developers to write their applications in popular languages ​​like Python, use these major AI frameworks, and their applications will virtually run on Nvidia chips.

Most of these application-level startups won’t have the resources or know-how to write their own kernels — the low-level software that runs the chips — and port their applications to other AI chips. So Infinity is trying to create an alternative CUDA kernel software that works with any type of chip, such as SRAM, GPUs, phone chips, and systolic arrays. Infinity is part of a new wave of startups that are trying, product by product, to chip away at Nvidia’s market dominance.

Infinity is trying to create a universal inference library that runs on all chips, allowing those chips to automate the replication of recent research results.

Infinity was launched last year by Jeremy Nixon, who was a Google Brain researcher and founder of the hacking community AGI House. Nixon told TechCrunch that he decided to launch this company because he was obsessed with the idea of ​​“automated invention” — the belief that “AI systems could actually be metatechnology.” He said that he himself invented a machine learning algorithm called Omega, which essentially created new machine learning algorithms and automatically evaluated them in a feedback loop.

This success got him thinking about other cases where this approach could work, so he turned to hardware, believing that automated systems could also generate the low-level code, such as kernels and so on, needed to help chips run more efficiently.

Infinity’s AI research agent Ignition aims to write the low-level code needed for AI inference on replacement Nvidia chips. It tests, debugs, measures how quickly devices perform with the code, and automatically rewrites the code if necessary to improve performance. The system is self-improving, meaning it is constantly learning and improving itself. Nixon says it also adapts to different chip architectures, regardless of special designs. The result is what Infinity claims is a CUDA-level software package.

Customers include AI chip maker (and would-be Nvidia competitor) D-Matrix, and Infinity is in talks with other large chip and cloud companies, Nixon said.

However, humans are in the loop, providing high-level guidance while the agent does more of the tedious, tedious work. in One case studyThe startup has found that the agent works much faster than a human alone, reducing what could take years or months to hours or days. Infinity does not charge a licensing fee upfront; Instead, it takes a portion of the performance gains and cost savings, measuring changes in tokens per second.

Currently, Infinity has 26 employees, including those in design, operations and engineering.

When you make a purchase through the links in our articles, We may earn a small commission. This does not affect our editorial independence.

Leave a Reply

Your email address will not be published. Required fields are marked *