Editor’s Note: Following article based on excerpts from an interview with Dr. Fei-Fei Li at AI Startup School in San Francisco. Dr. Fei-Fei Li is often called the godmother of AI—and for good reason. Before the world had AI as we know it, she was helping build the foundation. She is best known for establishing ImageNet, the dataset that enabled rapid advances in computer vision in the 2010s. She is the author of The Worlds I See.
My entire career is going after problems that are just so hard, bordering delusional. To me, AGI will not be complete without spatial intelligence and I want to solve that problem. I just love being an entrepreneur. These words, delivered with a combination of humility and audacious vision, set the tone for a conversation that traverses decades of AI research, from the frigid winters of sparse data to the roaring summers of generative models. When Dr. Fei-Fei Li speaks, one is reminded that the story of artificial intelligence is not merely about algorithms or datasets; it is about the audacity to see what others cannot, to pursue the invisible threads of intelligence that weave through the natural world.
She recalls her earliest forays into AI, a time when the field seemed almost unrecognisable to contemporary eyes. “I was a first-year assistant professor at Princeton. Oh wow, hi Tigers,” she said with a playful self-recognition of her beginnings. At that moment, “the world of AI and machine learning was so different… there was very little data. Algorithms, at least in computer vision, did not work. There was no industry. You know, as far as the public was concerned, the word AI didn’t exist.” And yet, even in that emptiness, there was a dream. The dream of making machines think, of imbuing them with vision, with understanding. “Visual intelligence is not just perceiving. It’s really understanding the world and doing things in the world. So I was obsessed with the problem of making machines see.”
This obsession culminated in ImageNet, a project that has become almost mythic in the annals of AI. Dr. Li recalls its origins with the clarity of someone who has both lived history and witnessed its ripple effects: “Fast forward around 2007-ish, my student and I decided that we have to take a bold bet. We have to bet that there needs to be a paradigm shift in machine learning, and that paradigm shift has to be led by data-driven methods. And there was no data. So we’re like, okay, let’s go to the internet, download a billion images, that’s the highest number we can get on the internet, and then just create the world’s, the entire world’s visual taxonomy. And we use that to train and benchmark machine learning algorithm. And that was why ImageNet was conceived and came to life.”
Yet even with the data, the journey to impact was not instantaneous. Dr. Li recalls a period of uncertainty, years in which the promise of ImageNet had not yet been fulfilled. “Between 2009, we published this tiny little CVPR poster in 2009 to 2012, the AlexNet. There were three years that we really believe that data will drive AI, but we had very little signal in terms of if that was working.” Their approach, radical for its time, was rooted in openness. “We open-sourced. We believed from the get-go we have to open-source this to the entire research community for everybody to work on this… and then the first couple of years was really setting the baseline. You know, the performance was in a 30% error rate. It wasn’t zero or, I mean, it wasn’t completely random, but it wasn’t that great.”
It was the advent of AlexNet in 2012 that transformed uncertainty into revolution. Dr. Li remembers the night with an almost cinematic intensity: “I was home and said, we got a result that really, really stand out and you should take a look. And we looked into it. It was convolutional neural network… There was a couple of tweaks in terms of the algorithm, but it was pretty surprising at the beginning for us to see that there was such a step change.” That moment marked a convergence of data, computation, and algorithmic insight, a triumvirate that would shape the trajectory of computer vision for a decade.
But Dr. Li’s vision did not stop at objects. The human eye, after all, does not see isolated items; it perceives the world as a story, a scene, a context. “I thought it was a 100-year dream, which is storytelling of the world… You actually see a conference room with screen, with stage, with people, with the crowd, the cameras. You actually can describe the entire scene. That’s a human ability that is at the foundation of visual intelligence.” She and her students pursued this audacious goal, developing models capable of image captioning, transforming still frames into narratives. “I almost felt like, what am I going to do with my life? That was my lifelong goal. It was such an incredible moment for both of us.”
This pursuit of storytelling in machines led naturally into the era of generative AI. A playful moment between mentor and student presaged the transformation of computer vision into creative intelligence: “I actually joked with him. I said, hey, Andre, why don’t we do the reverse? Take a sentence and generate an image. Of course, he knew I was joking. He said, ha-ha, I’m out of here. The world was just not ready.” Yet today, the world is ready, and the technology has evolved to make text-to-image generation almost commonplace.
Underlying these technical feats is a philosophical throughline: the ambition to capture spatial intelligence, to understand and model the three-dimensional world with the same fluency that humans perceive it. “Solving the problem of spatial intelligence, to understand the 3D world, to generate the 3D world, to reason about the 3D world, to do things in the 3D world is a fundamental problem of AI. To me, AGI will not be complete without spatial intelligence. And I want to solve that problem.” The complexity of this undertaking cannot be overstated. Unlike language, which is a purely generative signal, the world exists in dimensions, obeying physics, offering only partial projections to the observer. Vision is fundamentally ill-posed, a mathematical puzzle layered with uncertainty.
Dr. Li articulates this challenge with precision: “First of all, the real world is 3D. And if you add time, it’s 4D. But just let’s confine ourselves within space. It’s fundamentally 3D. So that by itself is a much more combinatorially harder problem. Second, the sensing, the reception of the visual world is a projection. Whether it’s your eye, your retina, or a camera, it’s always collapsing 3D to 2D… And third, the world is not purely generative… You are now suddenly dialing between generation and reconstruction in a very fluid way.”
Her own career mirrors this oscillation between creation and reconstruction, between theory and application, academia and entrepreneurship. “I just started a small company, so very excited to be here… My entire career is going after problems that are just so hard, bordering delusional. And I think this is this is the delusional problem.” With World Labs, Dr. Li and her team tackle what she calls “the hardest problem in AI right now,” assembling a cadre of prodigious talent to confront the challenge of modeling and generating 3D worlds. These world models promise utility across creation—from designers, architects, and game developers—to robotics, metaverse content, and beyond.
Yet Dr. Li’s story is also deeply human, shaped by resilience and the lived experience of starting anew. She recounts arriving in the United States, with no English and limited resources, running a laundromat to fund her studies, and later navigating the uncharted waters of academia: “I was 19 and I was out of desperation… I fundraised. I was the founder CEO. I was also the cashier and all the other things. And I exited. So after seven years.” This narrative of persistence and ingenuity underpins her professional ethos: “Forget about what you have done in the past. Forget about what others think of you. Just hunker down and build. That is my comfort zone. And I just love that.”
Mentorship, too, remains central to her identity. Advising students like Andrew McCarthy, Jim Fan, and Jia Deng, she has guided a generation of researchers who now shape the field she helped define. “I’m the lucky one… I could tell early on who was going to change the field of AI.” Her ability to see potential mirrors her approach to machines: a patient, deliberate cultivation of understanding, whether human or artificial.
Dr. Li’s journey is thus both personal and emblematic, bridging epochs of AI research, moments of technical breakthrough, and the intimate truths of human perseverance. From ImageNet to World Labs, from objects to scenes to generative 3D worlds, her work embodies a rare blend of scientific rigor and visionary daring. In her own words, “I really care about creating a beacon of light in the progress of AI and try to imagine how AI can be human centred, how we can create AI to help humanity.” And it is precisely this synthesis—of ambition, intellect, and empathy—that renders her a guiding figure not just in AI, but in the ongoing story of human imagination itself.

