Each use case has tens of funded companies. Each is churning out features rapidly, getting to parity faster than customers can imagine. Each has early traction and a worthy claim to win.
What will it take to eventually win the game?
1) Will it be about surviving the multiple shakeouts that each vertical/ use case will eventually see? Letting capital-bloated companies implode and letting the “tourists” give up…
2) Will continued product obsession be the key? Essentially refining the product beyond where others give up…
3) Will choosing non-obvious wedges/ ICPs be the way to differentiate & survive? Serve markets that others are choosing to ignore/ finding unviable to serve…
The technology is still so early, and we clearly have a few decades of upside left. Yet, there is a gold rush going on right now, which I am sure will push people to optimize for the short term.
In that case, will founders who are truly playing the long game ultimately win? Or is it more important to “surf the wave” in the present?
The former will look unattractive in current times and hence, will be undervalued and “contrarian”. The latter will appear to be imminent winners, yet could flame out.
Reporting back a few thoughts running through my head in real-time:
1. Essentially, the capabilities & design of every core SaaS use case are being reimagined by AI founders as we speak. In a future steady state, I see many of them living inside larger product suites as “features”, either via the incumbent fast-following and shipping them, or via small acquisitions/acqui-hires.
2. Consumer AI products remind me a lot of the 1st gen iPhone apps. Founders (developers) rapidly shipping entertaining, almost “toy-like” use cases. Like in mobile, will something massive eventually come out of these? So hard to tell…
3. An underlying capability of AI that a majority of products seem to be leveraging is “contextual artifact creation”. Eg. creating videos & decks in real time, replacing specific elements instantly in pre-existing media etc.
4. While the underlying “intelligence” capabilities of the products seem to be next-level, the UI/ UX as of now seems quite incremental relative to the mobile/cloud era. Lots more discovery & risk-taking needs to happen here.
5. Across enterprise & consumer/prosumer, it’s clear that these products can only manifest their power when they have access to extremely differentiated & diverse sources of data. In some contexts, it was unclear how a startup would get access to many such datasets in a fresh & relevant manner.
6. In legacy industries like govt/ public sector, AI-native products, even with game-changing capabilities, will still need to deal with age-old GTM challenges (long sales cycle, who will buy, what are the incentives for users to adopt etc).
7. Finally, it’s still pretty effin’ hard to pull off a glitch-free, low-latency AI demo.
Congrats to all the presenting SPC founders. Can’t wait for how these products shape up going forward!
Driven by an ambitious talent pool, geopolitical tailwinds, operating model innovation & domestic risk capital, Indian vertical AI startups could be the breakout tech story of the decade.
A potential scenario running in my head on how 🇮🇳 startups get a significant share in global AI over the next decade:
AI “infra” winners get built in 🇺🇸 (OpenAI, Alphabet etc.) ➡️
AI “platform” winners too, emerge in 🇺🇸 (Salesforce & HubSpot equivalents; a bunch get built/ led by the Indian diaspora) ➡️
As 1st-gen “Application” winners emerge in 🇺🇸, 🇮🇳 startups fast-follow in specific enterprise verticals & grab market share.
The time lag to fast-follow is significantly lower than, say, what Zoho did to Salesforce, or Freshdesk did to Zendesk.
This time, they play an asymmetric game. Instead of only competing with US startups on their home turf, Indian enterprise AI startups also look to dominate the Global South (SEA, MENA, LatAm, etc.).
Moving beyond binary operating models of India or US-based, Indian enterprise AI startups innovate & develop new, globally fungible, cross-geo operating models, similar to Infosys in the 90s, BPOs/ KPOs in the 2000s, and Chennai SaaS in the 2010s.
Compared to the SaaS wave, Indian enterprise AI startups get 10-100x more market share in each vertical, driven by a more ambitious & courageous founder pool, a talent base with skillsets & knowledge from previous tech waves, democratized knowledge & tools access courtesy of AI, as well as more availability of domestic risk capital at each stage.
Rather than IPO or M&A in the US, verticalized Indian enterprise AI startups either go public domestically, or get domestic Private Equity & conglomerates on the cap table who help them scale way beyond the last gen of software companies.
All these games play out on top of a favorable geo-political alignment between India & rest of the democratic world, driven by a China counter-balance narrative.
Verticalized Indian enterprise AI startups could be the contrarian venture bet of this decade!
During my recent India trip, a question I got asked repeatedly by both founders & investors was, “What are you seeing as the main differences between the AI ecosystem in the Valley vs India?”.
I currently see 2 main differences:
1/ Exposure (& therefore, Ambition)
AI founders in the Valley seem to have significantly more direct exposure to the work happening at the frontier. And not just in terms of the foundational technology, but also what battles the incumbents are taking on, how workflows are being iterated on, what lean, full-stack startup teams are doing to be able to generate significant product velocity & revenue, and how customers are thinking & allocating resources.
Essentially, they have the advantage of directly drinking from the Bay Area fountain of knowledge & information, spread primarily via networks.
A direct consequence of more exposure is that it uplevels the ambition of Valley AI founders and organically pushes them to raise the bar for execution within the company. Thus leading to sharper thinking, more courageous bets, and faster execution that all put together, improves the odds of a large outcome.
2/ Story-telling
I see that while AI founders in both the Valley and India are picking very similar problem statements to work on, the storytelling around the same use cases in the Valley is significantly superior.
I guess one reason is that operating directly in the target market (vs being a few degrees of freedom away from it) makes it much easier to get higher-quality early validation signals, making the story much more believable.
Also, AI founders in the Valley tend to emerge from the leading-edge companies of the last mobile/ cloud/ SaaS cycles. So they have a much better intuitive understanding of how to position & message the company in the early days to customers, investors & key hires.
Story-telling becomes even more important as how the AI landscape will evolve in specific market segments & verticals remains highly fuzzy.
So, what can India-based AI founders do to bridge these 2 gaps? Here are a few actionable things:
1/ Do extended sprints in the Bay Area regularly to drink from the same fountain.
2/ Surround yourself with Bay Area-based operators, angels & advisors (even remote is ok to begin with) who can regularly feed this knowledge & intel and, more importantly, help uplevel your thinking & ambition.
3/ Follow a conscious 0-to-1 strategy of only building for US design partners, so your product is held to the same bar as those from Valley startups.
4/ Specific suggestion for VCs – mine your network of LPs, Advisors & Portcos to hold regular AI knowledge sharing sessions with leaders of marquee AI-native companies that are building on the frontier in the Bay Area.
Given the exponential rate of change in AI, Dev Tools appear to have the most risk of use case durability, compared to Infra and Applications.
Earlier this week, I attended the AGI Builders Meetup in San Francisco. The event had 5 product demos ranging from new AI features by Twilio to a YC S24 startup building an AI agent to handle calls on behalf of users.
These demo-based events are always helpful for me to track the latest in AI. Also ended up chatting with a few engineers and founders attending the event, to get their thoughts on what they are seeing in their respective domains within AI.
One key insight I got from this event was that LLMs aren’t the real future of AI. No one really knows what’s going on inside them. They hallucinate much more than desirable (especially for accuracy-driven enterprise use cases). They are prone to prompt injection hacking.
In fact, the presenting Twilio PM said that “working with LLMs is like getting toddlers to do something”. They don’t take instructions promptly. You don’t know what’s going on inside their heads. You have to proactively ensure their safety as they are doing a task. I found this framing really interesting.
Building on this further, I have an updated (but still working) POV on the 3 buckets of AI – infra, applications, and dev tools.
#1 Infra
Irrespective of where AI goes, hardware infra like GPUs will always be needed. Hence, it makes sense that Nvidia is doing so much capex.
Also, Big Tech software infra players (Microsoft, Google, Meta, AWS, etc.), as well as AI-native hyperscalers (OpenAI, Anthropic, etc.), will continue playing a key role in defining where AI goes from here.
Sadly, as a micro VC, I can’t play much in this bucket (except personally investing in the public markets).
#2 Applications
Application layer founders that are starting up today are leveraging the capabilities of AI from Day 0 to solve customer problems. Their core focus still remains commercial-first – using the best-available software capabilities to solve customer problems, rather than getting overly enamored by the research aspects of AI and where it’s headed.
In a sense, these startups are centered on customer problems, not AI per se. Wherever AI ends up going, LLMs and beyond, these founders will leverage whatever capabilities they can get their hands on, and modify their architectures accordingly.
As an investor, the key is to back founders who are starting up now with an AI-first mindset and therefore, are fresh enough and agile enough to keep evolving their software as the underlying AI capabilities evolve.
Therefore, it’s reasonable to expect that these AI-native application startups should be fairly resilient to changes in the overall AI landscape. Hence, I feel reasonably comfortable in backing them (eg. portcos like Confido Health, Loop, and Soulside).
#3 Dev tools
This is the bucket I am most confused (and concerned) about. From seeing these demos, it seems like dev tools startups are essentially using the mental models of the previous cloud & mobile waves to make assumptions on use cases.
Further, I have observed that many of them are solving short-term, immediate pain points that could easily become irrelevant due to where AI goes from here, and/ or from the competition (eg. open source alternatives, AWS quickly launching it as a feature, etc.).
As I was seeing these demos, I looked up how much capital some of these companies had raised. Many of them have raised anywhere from $10-35Mn. The capitalization of these companies seems out of sync with the durability of their underlying use cases and revenue.
Essentially, what all this means is that I have a macro “Why Now?” question around the AI dev tools bucket. A top Bay Area engineer who recently left a cushy Big Tech job to start up was recently saying – “Given how things are changing every month, I am really not sure what to build right now”. I feel this is an intellectually honest view, rather than a FOMO-based approach that many VCs are taking.
Based on what AI practitioners like this person are saying about the exponential rate of change in AI, I fear that a majority of these dev tool use cases won’t endure.
Again, this is just my working POV. Would love to hear your views on what you are seeing.
Subscribe
to my weekly newsletter where in addition to my long-form posts, I will also share a weekly recap of all my social posts & writings, what I loved to read & watch that week + other useful insights & analysis exclusively for my subscribers.
The current AI landscape is giving strong mid-90s telecom boom vibes. Studying that cycle suggests that we might only be at the beginning stages of a multi-year bull cycle, wherein the ongoing massive investments in chips, infra and foundational models will ultimately enable the rise of enduring AI applications.
The last few days have been hectic in the world of AI. Introspecting on how things are unfolding at present in public, private, and commercial markets, where AI is right now is giving me major mid-90s telecom boom vibes.
Hyperscalers continue to raise huge $$$
1/ Microsoft did a sort of acqui-hire of Inflection AI for $650Mn via an interesting licensing deal structure. Given that Inflection was one of the high-flying foundational model startups and had raised $1.3Bn from Microsoft and Nvidia (cash + cloud credits) at a valuation of $4Bn in June last year, it’s unclear whether this is a good or bad outcome for employees and investors. Although, going by this tweet from Reid Hoffman where he mentions “good future upside”, looks like shareholders got some sort of equity package from this deal.
2/ Amazon concluded its initially committed $4Bn investment in Anthropic by investing the second tranche of $2.75Bn. Am assuming like other hyper scaler deals, this is a mix of cash and cloud credits. As per CNBC, this tranche was done at the first tranche’s valuation of $18.4Bn. The press release from Amazon hinted at deep integrations between both platforms, with Anthropic using AWS as its primary cloud provider internally for product development while also offering the latest generations of Claude-3 foundational models to AWS customers via Amazon Bedrock (a fully managed service for LLMs).
3/ As per The Information, Canadian pension fund PSP Investments is about to co-lead a fresh round of financing in Cohere at a ~$5Bn valuation. The company’s last round was at a ~$2.1Bn valuation last year, and given it’s reportedly at a $22Mn ARR currently, this is a rich revenue multiple to pay. Based on my conversations with BigTech AI operators, Cohere is significantly lagging Anthropic in terms of foundational model capabilities.
The current phase of AI seems to be like the beginning of the telecom boom in the mid-90s
Personally, I am finding it hard to predict which of the current foundational model hyperscalers and AI-first application companies will survive. Further, given the inflated valuations these deals are being done at, barring logo grabbing, I don’t see how investors can make outsized venture returns in these deals.
In parallel, while meeting super-early AI companies in categories like dev tools, security, and deep domain applications, I am struggling to see a clear right-to-win for a majority of them. Given data and distribution advantages of incumbent products both in Enterprise and Consumer, it’s unclear which seemingly-white spaces are actually viable startup opportunities.
However, amidst these struggles as a venture investor, I am feeling good about one hypothesis – all this capital going into infra and foundational model companies is actually building capacity for the next generation of enduring AI products to be built. This is quite similar to the role that in hindsight, the telecom boom of the 90s ended up playing for the adoption of the Internet.
As the Telecommunications Act of 1996 opened up the telecom sector to competition, a host of new entrants came in to become ISPs. They were followed by companies like Cisco, Ciena, Lucent, Nortel, and others who were desperate to sell networking equipment to these telecom companies.
Comparing this to today’s AI landscape, cloud providers seem to be similar to telecom companies, while semiconductor companies selling chips to these cloud providers are like the networking equipment companies eg. Cisco.
Also, during this boom, telecom capex was unlike ever seen before. As per the earlier cited post, just in the year 2000, capital spending by publicly traded telecom service providers was at an astonishing ~$120Bn (~$213Bn in today’s dollar terms).
This telecom boom capex is one of the largest capital bases ever built in such a short amount of time. I can see the same vibes in the amount of dollars going into AI chips, infra, and foundational models today.
Btw, one more learning from the telecom boom is how these flywheels become even stronger as the adoption of new tech starts reflecting in productivity gains. Here are some interesting excerpts on this from the Fabricated Knowledge post:
As this telecom boom was unfolding, LTCM blew up and therefore, the Fed ended up cutting rates in 1998 to avoid negative ripple effects. This is like adding tons of gasoline to a raging fire, ultimately leading up to a massive dotcom bubble.
If I play out the AI cycle like the 90s telecom boom, we might only be at the beginning stages of a multi-year bull cycle, similar to say 1995-96 (perhaps the launch of ChatGPT is similar to the Netscape IPO?).
Investments into the buildout of AI infra could run into trillions of dollars. In parallel, it seems the public markets have already started pricing in some of the future promises of AI. Going by the telecom boom, this pricing-in of future expectations could significantly accelerate for several years from hereon, driving stocks of both the telecom service equivalents (Cloud providers) as well as the networking equipment equivalents (Nvidia and perhaps any new entrants into chip manufacturing?).
That all this is happening in a higher-interest rate environment is a critical point. If for any reason (economic, geopolitical, or otherwise) the Fed starts cutting rates (which they are publicly saying they will), this could provide a major kicker into an already accelerating bull market.
So, we can reasonably posit an oncoming AI bull market for the next few (at least 3-5) years. Ultimately, like all bull markets, it will transform into a bubble, which will then peak and eventually crash. If you look at the Internet wave, the massive telecom capex of the 90s ultimately enabled the rise of enduring Web 1.0 companies like Google and Facebook, but only after the dotcom crash. Hence, I tweeted this yesterday:
What do you think?
Bonus Section: Commentary On Sequoia Capital’s AI Ascent 2024
As I was trying to make sense of recent AI funding developments, I chanced upon the just-released videos from Sequoia Capital’s AI Ascent 2024. I found these points from the keynote particularly interesting:
1/ If we draw parallels with the Cloud wave, in 2010, the entire global software TAM was ~$350Bn, of which Cloud was a tiny ~$6Bn sliver. Cut to 2023, the global TAM has grown to ~650Bn but more importantly, Cloud has grown to a ~$400Bn large piece of this pie (~40% CAGR over 15 years).
The starting pie for AI is not just software products, but also services that can be automated. So the hypothesis is that the starting pie for AI is ~$10Tn.
2/ The “Why Now” for AI is really strong, wherein a set of additive waves, starting from semiconductors in the 60s to Cloud and Mobile in the 2000s, has brought us to this stage. The ingredients to take AI from research to commercial applications are all there today.
Personally, I feel Sam Altman created the ‘iPod’ moment for AI by taking the power of AI to everyday users via a step-change ChatGPT product. In parallel, Jensen Huang should get shared credits for this catalytic moment given Nvidia’s rapid progress on giving chips more power at smaller form factors and hopefully over the next few years, making compute considerably cheaper and easier to access.
3/ In the last Cloud and Mobile transition waves, a host of new categories were created and new leaders were born in each of them. For AI, most of the major categories – (a) Infra, (b) Security, (c) Data, (d) Developer, and (e) Apps, are open right now. Hence, a massive opportunity for new category leaders to be created.
Interestingly, by depicting the white spaces this way, Sequoia also seems to hint that the Infra category is likely to be dominated by BigTech incumbents in chips and cloud. I am also reading the sub-text that Sequoia, in a way, views hyperscalers like Anthropic to be embedded within the existing cloud ecosystems (and hence, no separate logos depicted).
4/ Sequoia estimates that Generative AI companies are clocking in ~$3Bn in annual revenues in aggregate at present. As a comparison, SaaS took 10 years to get to this aggregate revenue scale as an industry, something that AI has achieved in almost the first year out of the gate.
5/ One of the early signals that AI is a real transformative wave is the sheer traction that the early products are getting across both Enterprise and Consumer.
6/ Over the last year, a majority of the capital has gone into the foundational model companies. In the Web 1.0 wave, the Application companies that came later in the cycle (eg. Google) captured the most value. The current uneven distribution of funding indicates that the Applications layer in AI hasn’t even gotten out of the stables yet.
7/ The usage numbers of AI-first products are still way behind incumbents. Eg. the median DAU/MAU ratio of AI-first products is a mere ~14%, compared to ~51% for incumbent products. This indicates that AI adoption is still in its infancy.
It’s encouraging to see that the ability of foundational models is on a continuous upward trend. At some point, this will translate into product capabilities that meet the expectations of users, which will then eventually reflect in better usage and retention numbers.
8/ I loved this slide that showed how when the iPhone was launched, the first generation of apps were either gimmicky or basic utilities. It wasn’t until a few years later that companies learned how to harness the capabilities of the iPhone to build enduring products.
Reasoning by analogy, we should expect that it will take a few more years (though perhaps a smaller number than previous waves?) for enduring AI applications to emerge.
9/ Sequoia is calling AI primarily a “productivity revolution”, similar to farm mechanization. At a macro level, this should bring down the costs of doing any task or delivering services, creating a strong deflationary force in areas like education and healthcare that have historically seen a perpetual rise in costs.
to my weekly newsletter where in addition to my long-form posts, I will also share a weekly recap of all my social posts & writings, what I loved to read & watch that week + other useful insights & analysis exclusively for my subscribers.
Sharing some observations and working hypothesis on Opportunist vs Believer founding teams in AI.
My biggest challenge as a venture investor in AI right now is figuring out which of the following 2 camps a particular founding team belongs to:
Opportunists – who are trying to leverage this moment in time when the market has massive curiosity about AI.
vs
Believers – who have high conviction, and are truly mission-driven about AI.
This is a critical evaluation point for these early AI deals. As previous super-cycles have shown us, a bubble-bursting trough in the space is inevitable in a few years (perhaps as soon as 3-5 years?). It will be brutal like previous resets – capital will get reallocated to the winners and dry up for the rest, exits will be on brutal terms, customers will tighten their belts, early-stage talent will flee and the general sentiment will turn from greed to fear.
In my experience, Opportunist founding teams are less likely to survive this trough. It will require grinding out on fumes and focusing on real customer problems vs vanity metrics and perpetual fundraising. It will need gut-wrenching decisions that sacrifice short-term gratification so that the long-term upside can be captured. It will require possibly resurrecting the company many times from the dead.
Being able to do all this requires extremely high conviction deep down in the gut. Founders who are Believers will have this conviction in their DNA, and when the cycle turns negative, this will become their competitive advantage.
Given this is turning out to be a key evaluation point for AI deals, have been thinking through what leading signals can be used to spot Believers with higher probability. Here are some working hypothesis thoughts on this:
[Disclaimer: am just thinking out loud here so please take this with a pinch of salt. This is nowhere near any gospel of truth, nor do I have significant experiential validation around these points given we are literally in the first wave of AI deals].
1/ Pre-ChatGPT AI builders – likely to have been working in AI much before ChatGPT was launched. They were most likely building with ML, NLP, and neural networks in a Big Tech team, a lab, a university, or some sort of R&D/ academic environment.
2/ Pre-AI domain experts – likely to have been working deeply in a specific domain/ industry/ sector/ function from pre-AI days and are now adopting LLMs to carry forward their domain work and solve customer problems that were previously unsolvable or unviable.
3/ Young tinkerers – likely to be fresh grads who started building AI-native products as a hobby during university, maybe as part of a side hustle, or even just out of intellectual curiosity. They would have likely built products and hacked a few early users even without “doing a startup”.
These are only some of the personas I have been thinking through. As I meet more teams, I will keep adding to this list.
If one looks at how the early days of Web 1.0 played out (eg. in eCommerce and Search), most first-movers ended up dying. The 2nd generation companies leveraged both the market that was created by the 1st gen, as well as learnings from their failures, to create new categories and emerge as viable businesses.
History doesn’t repeat exactly but often rhymes, thus requires being even more thoughtful about which companies to back in this 1st generation of AI. In my case, as a US-India corridor investor, there is an additional complexity to think through – how will AI companies being built out of India compete with those in Silicon Valley? Who is most likely to be stronger in which part of the AI stack?
With domestic data being of strategic importance to each country and the rise of country-specific models, is AI going to be an extension of the globally decentralized software product/ SaaS story of recent years? Or will there be opportunities in ring-fenced, domestic AI in each major geography?
These questions and unknowns are what make the present times in AI investing both interesting and challenging at the same time. To manage this context, I am trying to be open-minded, learn fast, and think from first principles as much as possible. But at the same time, balancing this default-optimism stance with being non-trigger-hungry, consciously thoughtful, and taking the time to build personal conviction on each opportunity.
to my weekly newsletter where in addition to my long-form posts, I will also share a weekly recap of all my social posts & writings, what I loved to read & watch that week + other useful insights & analysis exclusively for my subscribers.
Discussing two areas startups can focus on to create competitive differentiation in the AI applications layer – (1) data and (2) product flows.
Over the last few weeks, I have met a few exciting startups building in the applications layer of AI. As the landscape stands today, the foundational LLMs layer is likely to be dominated by a mix of open-source, Big Tech, and perhaps 1-2 hyper scalers (eg. Anthropic). The cloud infra (compute, safety, security etc.) to deliver these model capabilities will definitely be served by the Big Tech cloud players.
This leaves 2 categories for startups to exploit against these large competitors – (1) Applications and (2) Dev Tools. On the latter, I don’t understand it deeply enough to have a view on it (yet). However, the Application layer is something I get, and therefore, have some working POVs on it.
Almost all AI application layer startups I am seeing right now are essentially using ChatGPT (+ Bard and Llama in a few cases) to build features that solve sharp use cases in specific verticals. Based on observation, some low-hanging verticals that founders are going after include Insurance, Marketing, Sales, and HRTech with AI-generated content being a horizontal ingredient in most of these products (eg. automated email generation, stitching together a marketing video, crafting a training course outline etc.).
In all these cases, I am still struggling to understand how these startups can create competitive differentiation or moats purely by building features on top of hyper-scaler APIs. To take a step back:
For a new technology inflection to create viable startup opportunities, there need to be sizable areas where new companies are significantly better positioned than incumbents to leverage this new technology and solve unaddressed customer problems.
This is a really important point. For a startup to be viable, it’s not enough to just be an early adopter of cool technology and build new products before anyone else. The startup has to be able to create significant differentiation against entrenched competition too. Eg. Apple beat IBM in the PC inflection, Amazon beat offline retailers in the Internet inflection, Instagram and WhatsApp beat Facebook in the mobile inflection, and Figma beat AdobeXD in the cloud inflection.
This is the aspect where I am pushing all AI application founders I meet to start thinking through and strategizing from Day 0. A couple of ways to potentially drive competitive differentiation have emerged from these working sessions:
1/ Access to data
While ChatGPT is great for bootstrapping specific use cases, eventual product differentiation will emerge from startups fine-tuning their own LLMs (with open-source models as a starting point) using proprietary data sets for industry-specific use cases.
To put it simply, foundational models will keep doing a great job of adding horizontal knowledge. Startups will need to do the work of incorporating deep vertical knowledge into the models.
Here, access to the ‘right’ customer data will be critical. But then, entrenched incumbents would already have access to much more data than a 0-to-1 startup. So, how does a startup create a data advantage?
One way could be to identify unsolved pain points for customers that large pre-AI competitors aren’t going after, either because they are contextually unviable (Innovator’s Dilemma), were unsolvable pre-AI, or due to organizational inertia.
In these cases, AI-native startups can leverage their speed to get to the ‘right’ customer data sets before anyone else, and start creating an edge via custom fine-tuning and benefiting from faster learning cycles.
It’s interesting that the underlying driver of this differentiation is still good-old startup execution, rather than just building AI-first features. The company would still require classic software execution (founder-led sales, figuring out ICP, setting up GTM motions, etc.) to succeed.
2/ New product flows
Another area where startups could do better than the entrenched competition is putting in the work to develop AI-first product flows. We saw this happen in previous tech inflections where new capabilities and form factors gave rise to new ways of doing specific jobs. Eg. Apple cracked the smartphone user experience while Nokia struggled. Or Figma figured out how designers should work and collaborate with other functions in a fully hosted, in-browser experience, while Adobe continued to be stuck in its old UX.
Given the vast range of new capabilities that AI is unlocking (eg. chat-based UX, AI ‘agents’ to deliver specific tasks that underly use cases), it’s reasonable to expect a plethora of new workflows to emerge across customer segments. Many of them will require absolutely fresh product thinking to crack, something that pre-AI product teams at established companies might struggle with.
If one looks at both the above areas of potential startup differentiation, the way AI might end up creating viable startup opportunities is not the LLM technology itself, which will become baseline, widely available (like cloud today), and likely open source (similar to programming languages like Java and Python).
Rather, the drivers of value creation by startups will be in:
(1) What’s needed to effectively leverage these LLMs to solve verticalized, deep industry-specific problems – eg. pre-AI, a 10x backend engineer was needed to leverage the cloud. Post AI, specific datasets will be needed to leverage LLMs.
(2) 2nd and 3rd order impact of AI on product experiences and workflows – Figma and Notion took years of fresh thinking and iterations to reimagine collaboration UX in the cloud. AI-first use cases will require similar untethered, ground-up product thinking to deliver these capabilities effectively to customers.
What does this mean for venture investing?
It means even in the post-AI world, investors should continue to look for founding teams that demonstrate many of the classical startup traits, a few of them being – (1) ability to unearth a unique customer insight, (2) product thinking to be able to solve for it, (3) GTM skillset to be able to create a differentiated business out of it, and (4) grit to last through the journey.
Essentially, might be a good idea to avoid AI overthink and keep doing more of the basics of venture capital.
Note: check out the previous post #3 in this AI Musings series – LLMs for Beginners.
Subscribe
to my weekly newsletter where in addition to my long-form posts, I will also share a weekly recap of all my social posts & writings, what I loved to read & watch that week + other useful insights & analysis exclusively for my subscribers.
Sharing my notes from an awesome talk by Andrej Karpathy (top researcher at OpenAI) titled ‘Intro to Large Language Models’. A simple and quick primer on AI and LLMs for a general audience.
Just listened to this awesome presentation by Andrej Karpathy (top researcher at OpenAI) on The Busy Person’s Intro to LLMs. I found it extremely helpful and was making notes throughout. Sharing them below for anyone who wants to understand AI and LLMs in the simplest way possible:
1/ What is a Large Language Model (LLM)?
LLMs are intelligent pieces of code that can learn from a vast universe of data, make inferences using it, and use those inferences to answer a user’s questions or perform certain tasks.
An LLM typically consists of 2 files – (1) Parameters file and (2) Run file.
a) Parameters are the weights that go into the LLM and power the neural network within it.
b) The Run file is some sort of code to run logic on the neural network within the model.
LLMs can be of 2 types:
a) Proprietary – users don’t have access to the actual parameters/ weights within the model, and can only use it via a web interface or API. Eg. OpenAI’s ChatGPT, Google’s Bard etc.
b) Open source – users have access to the full model, including the parameters/ weights. They can easily modify the model for their own specific purpose. Eg. the Llama-2 model by Meta.
Here’s the example of Llama-2 given by Andrej, wherein the Parameters file is 140GB and the Run file is just 500 lines of code written in any language like C, Python etc.
Interestingly, you just need a standard computer to be able to run an LLM like Llama-2 and derive inferences. It runs locally on the machine and hence, even Internet connectivity isn’t a requirement.
However, training an LLM is a significantly more computationally heavy task that requires thousands of GPUs and needs millions of dollars in investment.
2/ What is the LLM’s neural network really doing?
Very simply, a neural network tries to predict the next word in a sequence with the highest confidence. Eg., for the sequence “cat sat on a…”, the next word is predicted to be “mat” with 97% confidence.
The next word prediction task forces the neural network to learn a lot about the world and assimilate tons of knowledge on the Internet, including the interconnections within it.
The outcome of this basic learning and word-prediction capability is the ability to create various kinds of documents eg. a piece of Java code, an Amazon product description, or a Wiki article.
The now-famous Transformer architecture is what powers this neural network.
3/ How the neural network “really” works is still a mystery
While we can input Billions of parameters into a network, and iteratively adjust them for better predictions, we really don’t know how the parameters are exactly collaborating within the network
We know that these Parameters build and maintain some kind of knowledge database, but it’s a bit strange and imperfect. For eg., the model can answer “who is Tom Cruise’s mother?” as “Mary Lee Pfeiffer” but if you ask “who is Mary Lee Pfeiffer’s son?”, it answers “I don’t know”. This is popularly called a Recursive Curse.
LLMs aren’t like a car, where we know each part and how they exactly work with each other.
Think of LLMs as mostly inscrutable artifacts. The only thing we truly know about them is whether they work or not, and with what probability.
We can give LLMs inputs, and empirically evaluate the quality of outputs. We can only observe their behavior.
4/ How do you obtain the secret sauce of an LLM – the Parameters?Stage 1: Pre-training
Getting the Parameters requires essentially “compressing the Internet”.
To illustrate how Llama-2 was trained, imagine taking a chunk of the Internet (~10TB of text, procured by crawling the Internet). You then compress this large chunk of Internet data dump using 6,000 GPUs* for 12 days. FYI, this would cost ~$2Mn as GPUs are super expensive. [*A graphics processing unit (GPU) is an electronic circuit that can perform mathematical calculations at high speed. This is what Nvidia makes and AI is the reason behind its stock ripping.]
What this GPU computation throws out is the Parameters file – a compression of all this Internet data, almost like an intelligent distillation of it.
5/ Stage 2 of training – Finetuning
The document generation capability of a base-trained LLM is of limited use. What we want is a Q&A-type assistant model. The way to achieve this is to take the model through the next stage of training called Finetuning.
In this stage, we swap out the generic, large Internet dataset, and replace it with a specific Q&A dataset that has been collected manually.
This step of finetuning is done by companies like OpenAI and Anthropic, who hire people that create ideal answers to questions that users might typically ask. This human-curated dataset of conversations then becomes labeling instructions for the LLM to be finetuned upon.
It’s important to note that at this stage, quality is more important than quantity.
After finetuning, your base LLM becomes an assistant model. Interestingly, and again we don’t understand how it really happens, but the model is able to use the generic Internet knowledge from the pre-training stage and combine it with the conversational datasets from the finetuning stage.
The below slide summarizes both parts of LLM training. Pretraining is very expensive so update training typically happens once a year. Finetuning is more of a model alignment exercise and can happen as frequently as once a week.
6/ Stage 3 of training – comparisons
There is another subsequent stage of finetuning that can happen where for each question, human labelers compare multiple answers from the assistant model and label the best one. OpenAI calls this step Reinforced Learning from Human Feedback (RLHF).
7/ How do labeling instructions given to humans look?
Here is an excerpt from InstructGPT paper by OpenAI. It shows how human trainers should craft responses that are helpful, truthful, and harmless. These instructions can run into hundreds of pages.
Rather than being totally manual, labeling is fast becoming a human-machine collaboration exercise where LLMs can create drafts for humans to splice and put together, LLMs can review and critique labels based on instructions etc.
8/ How to compare various LLMs and find the best one for your use case?
There are platforms like Chatbot Arena (managed by Berkeley) where models are compared against each other and ranked as per an Elo rating, kind of like comparing and ranking chess players.
As seen below, closed proprietary models are ranked much higher than open-source models. However, the former can’t be modified for your use while the latter can.
9/ LLM scaling laws
Performance of LLMs is a smooth, well-behaved, predictable function of (1) N – no. of parameters in the network and (2) D – the amount of text we train on.
And the trends don’t show signs of topping out.
We can expect “more intelligence for free” by scaling.
Essentially, algorithmic progress isn’t necessary for better models. We can just throw more N and D via more compute power and get better intelligence.
Another interesting point – we can get better ‘general capability’ simultaneouslyacross all areas of knowledge by training a bigger model for longer. Essentially, with more training, we should expect the performance of models to rise more across all areas for free.
This is the reason driving the gold rush for (1) data and (2) compute, as the more we throw of it at LLMs, we can get exponentially better models. Algorithmic progress then just becomes a bonus.
10/ Future Idea #1: System 1 vs System 2 thinking by LLMs
Currently, LLMs are System 1 thinkers, similar to what Daniel Kahneman calls the intuitive, fast-thinking part of the human brain (like speed chess). Where researchers are trying to go is converting them into System 2 thinkers, mimicking the slow, deliberate, rational, reflective part of the human brain (like professional chess players thinking through decision trees on the spot).
Turning LLMs into System 2 thinkers is about trading off time for accuracy.
11/ Future Idea #2: self-improvement
Another area of future improvement is making LLMs capable of self-improvement. The AlphaGo program by DeepMind has already demonstrated that in a narrow, sandbox environment with a clear reward function (telling the program whether how it played is good or bad), AI can self-improve to become better than the best human player. Can this be achieved in other, more open contexts?
12/ Future Idea #3: custom LLMs
This is already being done by OpenAI via GPT agents that can do custom tasks and are available on the GPT Store (more details in my post on the OpenAI Dev Day).
Bringing it all together – LLM OS
I don’t think it’s correct to think of LLMs as just chatbots or word predictors. Think of it as the kernel process of an emerging operating system.
This process is coordinating a lot of resources like memory and computational tools, for problem solving.
As per Andrej, this is how the LLM OS framework will likely look:
There is another similarity between today’s OS ecosystem and the LLM OS view. Today, there are proprietary OS like Windows, iOS, and Android, with a fragmented ecosystem of open-source products co-existing. Similarly in LLM OS, there are proprietary LLMs like ChatGPT, Claude, and Bard, with a burgeoning parallel ecosystem of open-source LLMs, many of which are presently powered by Meta’s Llama series.
Appendix: LLM Security
As per Andrej, one of the major security threats facing LLMs is various ways to jailbreak them and get answers to undesirable questions. For eg. a user can role-play within a question to make it sound harmless to the LLM. In other cases, users can create encoded versions of harmful questions and bypass the LLM’s security check.
Another type of security threat is prompt injection, where bad actors can infect images, documents, and web pages with harmful prompts that can cause undesirable behavior by LLMs when the user tries to input these materials.
The final type of threat is data poisoning where attackers can inject trigger words within training data that corrupts the model and catalyzes specific undesirable behavior.
Subscribe
to my weekly newsletter where in addition to my long-form posts, I will also share a weekly recap of all my social posts & writings, what I loved to read & watch that week + other useful insights & analysis exclusively for my subscribers.
1/ Amazing GPT-4/ Turbo upgrades and new features announced. In particular, loved the ability to upload docs into ChatGPT. Also, the ability to choose pre-programmed voice modalities that sound significantly more realistic than any current digital alternatives.
Was also awesome to see Coke’s campaign that lets its customers programmatically create Diwali cards using DALL.E 3.
2/ The icing on the cake was the introduction of ‘GPTs’ or agents. Users can now build AI agents within ChatGPT that absorb a set of instructions and then take specific actions while leveraging the GPT-4 expanded knowledge base.
3/ Building GPT agents in natural language is the democratizing aspect of Generative AI and something that was missing in the earlier voice-to-action apps/ personal assistants in the mobile paradigm.
Sam’s natural language demo reminds me of all the bottlenecks we faced while building first-generation mobile search/ deep-linking at Quixey + all the work my friend, the late Rajat Mukherjee, did on voice-to-actions at Aiqudo. AI is on track to solve all those engg./ product challenges.
4/ OpenAI also showcased the GPT Store, which will feature the best GPTs built by developers on a revenue share model. This AI app store is a natural extension of the democratized-agent strategy.
5/ The developer Playground demo was really interesting, demonstrating capabilities like threading, function calling etc.
Essentially, any developer can now build agents within their app for their customers. These agents can have all advanced GPT-4 capabilities that power specific use cases like trip planning, navigation, splitting expenses etc., each of which is presently done by separate siloed apps.
Was awesome to see the demo agent communicate in a Jarvis-like voice modality.
6/ Finally, stoked to see the love Satya Nadella showed OpenAI and Sam during a friendly on-stage banter.
It looks like the OpenAI partnership has given a new lease of life to Azure and maybe even a game-changing competitive advantage against other cloud providers. In parallel to all the work that OpenAI is doing on the model side, Azure is building a new end-to-end, AI-native cloud infrastructure and compute stack to support the development and GTM of these efforts.
It was also heartening to see Satya underline security as one of the core focus areas for the partnership:
We are grounded in the fact that safety matters. Safety is not something you care about later but it’s something we do shift-left on.
Satya Nadella at OpenAI DevDay
My TLDR take:
The rollouts in this first-ever DevDay by OpenAI are clearly important milestones in this rapidly evolving space. AI is becoming easier to use, more powerful, and more accessible at an exponential pace. Personally, this is the first time I am seeing a potential v0.1 of what has been a larger-than-life but fuzzy vision of AGI.
Kudos to Satya and Microsoft for what’s turning out to be a generational business bet on OpenAI that frankly, seems to have caught the other Big Techs a bit flat-footed. However, expect strong responses from Google, Meta, and AWS in the coming months.
Finally, I have met many founders over the last few months who have been building nifty micro-products on top of OpenAI. A few of them have been touting how these are large, venture-returns opportunities. This DevDay has already shown how many of these startup ideas have already become point features within the OpenAI ecosystem.
This aggressive feature rollout by OpenAI once again brings to the fore strategic questions around moats, right-to-win, feature vs product vs platform, and access to 1st party training data. All this is significant food for thought both for founders and VCs.
As Big Tech, OpenAI, and other hyper-scalers like Anthropic continue to dominate the infra and model layers, for new startups, things like sharp domain expertise, deep understanding of specific customer problems, access to proprietary 1st party data as well as industry or audience-specific distribution channels, can become important sources of sustainable competitive advantage and drive a valid case for why a startup should exist.
to my weekly newsletter where in addition to my long-form posts, I will also share a weekly recap of all my social posts & writings, what I loved to read & watch that week + other useful insights & analysis exclusively for my subscribers.