Friday, October 9, 2026 AI news, turned into opportunities PDF newsletter
AI Opportunity Daily
Subscribe
AI agentsVerificationAI securityEnterprise AI

Google's AI coworkers, OpenAI's 372 math claims, and AI tools hacked at Pwn2Own

Google wants one Gemini agent, with its own inbox and calendar, to take on whole assignments at work. OpenAI has published hundreds of AI-made math results that nobody has fully checked. At Pwn2Own, hackers broke into LiteLLM, OpenAI Codex and Oracle's AI database. All three stories point the same way: AI can now produce more work than people can check, so whoever does the checking, governing and securing has the advantage.

Key takeaways

  • Google's new Gemini agent can work as a 'digital coworker' with its own Workspace identity and can switch between Gemini and Claude models, but Google hasn't announced pricing or a general-availability date.
  • OpenAI released 372 families of math results from an unreleased model, including a claimed Unique Games Conjecture proof, but only part of them have computer-checked proofs and outsiders can't yet reproduce the work.
  • Pwn2Own Ireland's new AI Infrastructure category produced working exploits against LiteLLM, OpenAI Codex and Oracle's Autonomous AI Database on day one, which starts a 90-day patch window for teams that run these tools.
Story 1 of 3Tools6 min read

Google Cloud unveils a 'universal' Gemini agent that can work as an AI coworker with its own Gmail, calendar and Drive

Google's new enterprise agent takes a goal, runs for hours or days, coordinates sub-agents and can join a team with its own Workspace account. It can already route work to Anthropic's Claude as well as Gemini. Price and launch date have not been announced.

Google Cloud used its Gemini at Work 2026 event on Thursday, October 8, to announce the 'Gemini agent'. CEO Thomas Kurian called it a single, universal agent for work. Instead of offering a set of separate assistants, Google wants employees to give one agent a goal, such as researching a market, building a model from company data or drafting a deck. The agent then returns finished work inside Gmail, Docs, Sheets, Slides, Chat, Calendar or a developer environment. Google sums this up as giving it objectives, not instructions. The agent is in private preview. According to 9to5Google, Google says it will reach more Workspace customers on select Business and Enterprise plans soon.

The most unusual feature is what Google calls a 'digital coworker'. VentureBeat gives an example: a manager could create an 'Event Planner' agent with its own Workspace account, email address, calendar, Drive storage and company-directory listing. Colleagues would give it work by @-mentioning it in Google Chat or asking it to edit documents. Google says these coworker agents can only see what team members explicitly share with them. They do not get a human employee's full permissions. Each agent has a cryptographically attested identity, its actions go into audit logs, and an 'Agent Gateway' applies company-wide rules to how agents connect with outside systems.

The agent runs in Google's cloud, so a job can last hours or days and keep going after the user logs off. Google describes four kinds of memory: session (the current task), semantic (accumulated knowledge), procedural (how jobs get done) and episodic (past work). The aim is for the agent to learn a company's processes over time. It can set up temporary teams of sub-agents and react to schedules or events. It connects to tools including Salesforce, ServiceNow, Jira, Git, BigQuery, Snowflake, Databricks and Microsoft Teams, as well as any Model Context Protocol (MCP) server. An API also lets it run 'headless', inside other apps, with no interface of its own.

The agent is model-neutral. A 'Smart Routing' feature picks a model for each job based on performance and cost. For now it chooses between Google's Gemini models and Anthropic's Claude models, and Google says other proprietary and open models will follow. Google also announced versions for specific industries. Financial Services, in preview, includes more than 50 foundational skills and data from FactSet, LSEG and SEC filings. Legal, also in preview, is built around document-management permissions and supports iManage and NetDocuments. Versions for government, healthcare and retail are coming later. Google says nearly 90% of Fortune 100 companies use Gemini Enterprise, but that figure describes the existing product, not the new agent.

The launch comes during a crowded few weeks. VentureBeat lists Microsoft's updated Copilot with a background 'Autopilot' agent (September 25), OpenAI's always-on Dots agents (September 29), Anthropic's native Claude integration with Google Docs, Sheets and Slides (October 6) and SpaceXAI's persistent Grok Bot and Team Bots. All of these companies are moving toward the same idea: named AI workers that keep running and have their own memory and identity. What sets Google apart is that customers can keep the same agent setup, company context and skills while swapping the model underneath, even for a competitor's model. VentureBeat notes that Microsoft also supports multiple models.

Several important questions are still open. Google hasn't announced a separate price for the agent or said whether it comes with existing Gemini Enterprise subscriptions. It hasn't given a general-availability date or explained how coworker agents will be set up and licensed. VentureBeat also notes that Google hasn't published independent benchmarks showing how reliably the agent finishes long tasks that span several apps. Some of the cost controls are not new either. The project-level spending caps that pause an agent when its budget runs out were announced in August. Until pricing and reliability data appear, buyers can't judge what a 'digital workforce' would really cost to run.

For buyers, the choice is no longer just which chatbot to license. It is which platform they trust to hold company context, agent identities and permissions. For builders, the support for MCP and the headless API means a skill or connector written once can work inside Google's agent and in rival products. The hard problems are shifting to governance: what each agent is, what it can see, who approves its actions and who pays for its compute. Companies that set clear rules for this early will be able to adopt AI agents more quickly and with fewer mistakes.

Why it mattersThe fight over enterprise AI is shifting from whose model is best to whose platform holds company context, agent identities and permissions. Google is betting that being open to rival models, including Claude, will win buyers.

👀 What to watch

Watch for pricing and licensing details, a general-availability date, and the first independent reports on how reliably coworker agents finish multi-day tasks.

3 opportunities from this story

1

AI coworker governance setup for Workspace companies

Medium⏱ 3-6 weeks💰 Fixed-fee setup package plus a monthly retainer for access reviews and audit-log checks

Offer a fixed-price package that gets mid-sized Google Workspace companies ready to deploy Gemini coworker agents safely. The package would cover identity naming conventions, which data each agent can see, approval rules, audit-log reviews and budget caps. Google says coworker agents only see what people explicitly share with them, so someone has to design that sharing model. IT teams at 50-1,000-person firms rarely have time to do this, and the preview period gives you a few months to build expertise before general availability.

Best for
Google Workspace consultants, IT managed-service providers, freelance security and compliance advisers
First step this week
Write a two-page 'AI coworker policy' template covering agent identities, data sharing, approvals and spend caps, and send it to five existing Workspace clients as a free resource.
Open the full playbook
Launch steps
  1. Study Google's published material on agent identities, Agent Gateway, audit logs and budget caps
  2. Build a policy template and a checklist for reviewing permissions
  3. Pilot it with one or two friendly clients, ideally ones in the private preview
  4. Package the offer with a clear scope and price
  5. Add a quarterly access-review retainer
Tools
Google Workspace Admin consoleGemini EnterpriseGoogle Cloud audit loggingNotion or Google Docs for policy templates
Risks

Google hasn't published full provisioning and licensing details, so some of your work may change at general availability. Keep templates modular and sell the governance thinking, not settings for one specific feature.

2

Vertical MCP connectors and skills for niche business software

Medium⏱ 4-8 weeks💰 Per-company monthly subscription for a hosted connector, plus paid custom skills

Build and sell MCP servers that connect niche industry systems, such as practice-management, property-management or field-service software, to the Gemini agent. Because the agent accepts any MCP server, and Claude and other agent platforms support MCP too, one connector can serve several platforms. Google's industry agents cover finance, legal, government, healthcare and retail, which leaves many smaller sectors without integrations.

Best for
Indie developers and small dev shops with experience in a particular industry
First step this week
Pick one niche system with a documented API that your contacts use, and build a read-only MCP server with three useful tools this week.
Open the full playbook
Launch steps
  1. Interview three to five operators in the niche about the tasks they repeat most
  2. Build a read-only MCP server first, then add carefully scoped write actions
  3. Test it with Gemini Enterprise and at least one other MCP-capable agent
  4. Publish documentation and a security note on data handling and permissions
  5. Sell through industry communities and software vendor partner programs
Tools
Model Context Protocol SDKsGemini EnterpriseClaudeCloudflare Workers or Google Cloud Run for hosting
Risks

The software vendor may ship its own connector, or the platforms may change their APIs. Choose systems whose vendors are slow to integrate, keep the connector portable across platforms, and get the vendor's API terms in writing.

3

Agent platform bake-off service for enterprise buyers

Medium⏱ 2-4 weeks💰 Fixed-fee evaluation projects, with optional follow-on implementation work

Google, Microsoft, OpenAI, Anthropic and SpaceXAI all launched persistent work agents within about two weeks, and none has published independent reliability benchmarks. Offer a paid evaluation that runs a client's own 10-20 real workflows through two or three platforms. Report on completion rate, errors, human time saved and compute cost, using Google's per-project cost tracking where it is available. Mid-to-large companies need this evidence before committing budget.

Best for
AI consultants, operations analysts, evaluation and QA specialists
First step this week
Design a standard scoring sheet for evaluating workflows (success, number of corrections, time taken, cost, security incidents) and run it on one internal workflow in two agent products.
Open the full playbook
Launch steps
  1. Create a scoring method and a report template
  2. Get access to previews or trials of two or more agent platforms
  3. Run a pilot evaluation on your own workflows and publish an anonymized sample report
  4. Sell to operations and IT leaders who are drafting 2027 AI budgets
  5. Update results as platforms ship new versions
Tools
Gemini EnterpriseMicrosoft CopilotChatGPT EnterpriseClaudeGoogle Sheets or Airtable for scoring
Risks

Platforms change quickly, so results go stale fast, and preview access may be limited. Date-stamp every report, sell re-tests as a subscription, and stay vendor-neutral to keep your credibility.

Story 2 of 3Research6 min read

OpenAI posts 372 families of AI-generated math results, including a claimed Unique Games proof, and mathematicians ask to see the evidence

An unreleased OpenAI model produced hundreds of manuscripts that OpenAI says resolve or advance open problems, most of them, the company says, from a single prompt. Only part of the work has computer-checked proofs, outsiders can't reproduce it, and experts disagree about what it means.

At 6pm EDT on October 6, OpenAI posted a large set of mathematical results from an unreleased internal model to GitHub. Scientific American, which first reported the release, says it contains 372 result 'families'. OpenAI says each one resolves or substantially advances an open question in mathematics or theoretical computer science. A family groups a main theorem with related proofs and consequences, so there are more manuscripts than families. There were 722 at launch. According to Implicator, the count fell to 719 on October 7 after OpenAI withdrew three manuscripts because of a sign error and revised 14 others.

The claims are bold. They include a proof of the Unique Games Conjecture. This central open problem concerns how hard it is to find near-optimal answers to optimization problems such as Max Cut. The list also includes L = BPL in complexity theory, faster integer multiplication, a four-dimensional Kakeya result, progress related to the Riemann hypothesis, and a positive answer to the unitary synthesis problem that Scott Aaronson and Greg Kuperberg posed in 2007. On October 7, Aaronson wrote that no human appeared to have understood just about any of the proofs yet. An OpenAI spokesperson said many of them are not yet understood by OpenAI's own mathematicians either.

How the results were produced may matter even more. An OpenAI spokesperson told Scientific American that almost every result came from a single prompt given to a single AI agent, although some took several attempts. Decrypt reports that OpenAI says it gave the model about 4,000 problems and kept the outputs it judged significant. Each published result used about three hours of ChatGPT Pro thinking compute on average. OpenAI's September Navier-Stokes claim, by contrast, used a swarm of 10,000 agents and millions of dollars in compute. Because the model hasn't been released, nobody outside OpenAI can test the single-prompt claim.

Checking the work is the central issue, and sources disagree about how much of it has been machine-verified. Decrypt, citing the repository's own formalization catalog, says only 162 of 722 manuscripts have a main result formalized in Lean, about 22%. Lean is a programming language in which a computer checks every step of a proof. Implicator says that as of October 7, 300 of 719 top-line results, about 42%, had Lean proofs. The gap may come from fast-changing updates or from counting things differently. OpenAI itself warns that some unformalized results could have issues. Even a passing Lean check only shows that the formal statement holds. It doesn't show that the statement matches the original problem or that the result is new or important.

Mathematicians are divided. MIT's Andrew Sutherland told Scientific American that claims of one-shot solutions by a single agent should be treated as unverified until the model is released and others can replicate the results. On September 29, an advisory group at the Institute for Advanced Study asked labs to publish the model name, prompts, a reasoning summary, time and compute for every result. According to Decrypt and Implicator, OpenAI released averages and 10 reasoning summaries but no prompts. Engadget, by contrast, reported that OpenAI said it followed the group's guidelines. The Association for Human Mathematics urged mathematicians to stop working with OpenAI, while number theorist Abhishek Saha called it a very big day for mathematics.

There are practical problems too. Dana Moshkovitz, a UT Austin complexity theorist who spent years working toward the Unique Games Conjecture, said in texts Aaronson published that she could only read the claimed proof with help from AI. Decrypt notes that the repository has GitHub Issues turned off and had never accepted a pull request, so outsiders have no obvious way to submit corrections or new formalizations. Some people also worried that new mathematics could weaken encryption. Implicator reports that researcher Justin Drake described a worst-case scenario for ECDSA. OpenAI reported no attack on ECDSA or RSA, and Vitalik Buterin advised against moving wallets in a hurry.

If even some of these results hold up, the bottleneck in research moves from producing proofs to checking, understanding and crediting them. The same pattern is likely to reach law, engineering, software and science, wherever AI can produce output faster than experts can review it. Formal verification tools like Lean, and the people who can use them, become much more valuable. So do clear rules about disclosure, reproducibility and who is responsible for AI-generated claims. For now, treat the headline results as claims under active review, not as established mathematics.

Why it mattersThis would be the largest batch of AI-generated research claims so far, which makes verification, not generation, the scarce resource. How the math community handles it will set norms for AI-produced work in other fields.

👀 What to watch

Watch how much of the release gets Lean-checked over the coming weeks, what experts conclude about the Unique Games proof, and whether OpenAI releases the model or the prompts.

3 opportunities from this story

1

Lean formalization and proof-verification services

High⏱ 4-12 weeks💰 Contract work paid per project or per hour, research grants, or sponsored formalization bounties

Offer paid help turning human-written or AI-generated proofs into Lean, or checking existing formalizations against the original problem statements. Hundreds of AI-generated manuscripts have no formal check, and Lean only confirms the formal statement, so people who can link the formal version to the real question are in short supply. Possible buyers include university groups, journals, AI labs and grant-funded projects that need credible verification.

Best for
Math PhD students and postdocs, Lean/Mathlib contributors, formal-methods engineers
First step this week
Choose one unformalized result family from OpenAI's repository in your area, formalize a key lemma, and publish the code and a short write-up as a portfolio piece.
Open the full playbook
Launch steps
  1. Build public credibility with Mathlib contributions or a formalization from the release
  2. Write a clear service description: formalization, statement-fidelity audit, written report
  3. Contact journal editors, research groups and AI-for-math teams
  4. Price per project, with a small paid scoping phase first
  5. Partner with other formalizers for larger jobs
Tools
Lean 4MathlibGitHubLLM coding assistants for routine proof steps
Risks

Formalization is slow and hard to estimate, and AI tools may automate parts of it. Charge for scoping first, focus on checking that formal statements match the original problems, which is harder to automate, and get credit and authorship terms in writing.

2

A reading tool that maps dense AI-written proofs

Medium⏱ 6-10 weeks💰 Freemium web app with paid tiers for research groups and institutional licenses

Build a tool that takes a long AI-generated paper and produces a navigable dependency map. It would show each claim, where it is used and proved, whether it has a Lean check, and plain-language summaries. The expert who worked on Unique Games for years said she needed AI help just to read the claimed proof. Researchers, referees and technical due-diligence teams in other fields face the same flood of hard-to-read machine output.

Best for
Developers with NLP or LaTeX-parsing skills, academic software builders
First step this week
Write a script that parses LaTeX from one manuscript in OpenAI's repository into a graph of lemmas and their dependencies, and show it to three researchers for feedback.
Open the full playbook
Launch steps
  1. Parse LaTeX theorem environments and cross-references into a dependency graph
  2. Add LLM-generated summaries, with links back to the exact source lines
  3. Show which results have Lean checks wherever a formalization exists
  4. Pilot it with a seminar group reading the release
  5. Extend it to other long technical documents such as specifications or audits
Tools
Python with a LaTeX parserGraph visualization (e.g. Cytoscape.js)An LLM APIGitHub
Risks

LLM summaries can misstate the logic, which is dangerous in proof review. Link every summary to its source text, label summaries as aids rather than checks, and never present them as verification.

3

Crypto-agility inventory for fintech and crypto firms

Medium⏱ 2-4 weeks💰 Fixed-fee cryptographic inventory and migration-readiness assessments

The release revived worries that new mathematics could weaken widely used cryptography such as ECDSA. Help fintech, crypto and SaaS companies list where they rely on specific algorithms and how quickly they could switch, including to post-quantum options, without overselling the threat. OpenAI reported no attack on ECDSA or RSA, but boards are asking questions now, and a sober inventory is useful whatever happens.

Best for
Security consultants, applied cryptographers, fractional CISOs
First step this week
Write a one-page, hype-free briefing on what the release does and doesn't imply for ECDSA and RSA, and share it with five fintech or crypto security leads along with an offer for an inventory call.
Open the full playbook
Launch steps
  1. Build a checklist of where algorithms are used: TLS, signing, key storage, wallets, third-party SDKs
  2. Use scanning tools to find crypto libraries and certificates in use
  3. Score each system on how easily its algorithm can be swapped
  4. Deliver a prioritized migration roadmap
  5. Offer quarterly threat updates as a retainer
Tools
Certificate and dependency scannersOpen-source PQC libraries (e.g. liboqs)SBOM toolsSpreadsheet or GRC platform
Risks

Selling on fear of an unconfirmed break would be misleading and could backfire. Frame the work as standard crypto-agility hygiene that is worth doing anyway, and cite only verified developments.

Story 3 of 3Security5 min read

Pwn2Own Ireland's first AI Infrastructure category: hackers break LiteLLM, OpenAI Codex and Oracle's AI database on day one

On day one of Pwn2Own Ireland, researchers demonstrated 32 zero-days, including $40,000 exploits against the LiteLLM gateway and OpenAI's Codex cloud coding agent. Vendors now have 90 days to patch before details become public.

At Pwn2Own Ireland 2026, the Zero Day Initiative's hacking contest that opened in Cork on October 6, AI developer tools were publicly hacked for the first time in a dedicated category. According to ZDI, 21 entries on day one demonstrated 32 unique zero-day vulnerabilities. Targets included the Samsung Galaxy S26, the Philips Hue Bridge Pro, printers, the Sonos Era 300 and a new AI Infrastructure category. ZDI's figures, as reported by Daily Security Review, put day-one payouts at $388,500. Infosecurity Magazine reported 'over $368,000'. The difference is probably due to when each outlet tallied confirmed results.

Three AI-related products were hacked on day one. ZDI says Taisic Yun of Xint combined an improper input validation bug with code injection to get a reverse shell, meaning remote control of the server, on LiteLLM. LiteLLM is a popular open-source gateway that sits between applications and many model providers, and the exploit won $40,000. A second team, Out of Bounds, also exploited LiteLLM using four bugs. Because two of those bugs were already known, the entry was ruled a partial 'collision' and earned $15,000. Ikotas Labs needed only a single argument-injection bug to exploit OpenAI's Codex cloud coding agent, also for $40,000. Tech Insider reports that VinSOC chained five bugs against Oracle's Autonomous AI Database for another $40,000.

According to Tech Insider, the new category also included Chroma, an open-source vector database, and Nvidia Dynamo, an inference-serving framework. The outlet reports AI Infrastructure payouts of $135,000 on day one and $80,500 on day two. It puts total contest payouts at more than $608,500 after two days, compared with $792,750 at the same point in 2025. Its explanation is that the new AI category spread researchers' attention across more targets. The day-two figures come from one outlet and were described as pending a final tally, so treat them as provisional until ZDI publishes official totals.

The process from here is standard. Pwn2Own gives vendors 90 days to patch the flaws demonstrated at the contest before details are published. Technical write-ups of the LiteLLM, Codex and Oracle bugs are therefore not public yet, and defenders can't know exactly which versions or settings are exposed. OpenAI can fix Codex centrally because it is a hosted service. LiteLLM and Chroma are self-hosted by many teams, who will have to upgrade on their own once fixes ship. That step is often slow, which leaves systems exposed long after a patch exists.

The choice of targets is telling. An AI gateway like LiteLLM usually stores API keys for several model providers and sees every prompt and response that passes through it, so taking over the server can expose all of that. Coding agents like Codex run commands and handle source code. Argument injection tricks a tool into treating attacker-supplied text as command options, and it is exactly the kind of bug that can turn a helpful assistant into a way in for attackers. Vector databases such as Chroma often hold internal documents used for retrieval, which makes them valuable targets too.

Some caution is needed. Pwn2Own exploits are shown under contest rules against specific product versions, and none of the sources reviewed say these bugs are being exploited in the wild. The sources also don't include vendor statements on severity or patch timelines. Security vendor Aviatrix says the results show enterprise AI platforms need stronger network segmentation and anomaly detection. That is a reasonable conclusion, but it comes from a company that sells such services. Public details on each bug will only come after patches ship or the 90-day window ends.

The bigger signal is that the software running AI systems is now a mainstream attack target. A top hacking contest is paying for it at the same level as phones and routers. Over the past two years, many teams set up gateways, vector stores and agents quickly and skipped the hardening they would apply to a database or web server. Expect more AI categories at contests like this, more CVEs (public vulnerability IDs) in AI tooling, and more customers and auditors asking for proof that AI infrastructure is patched, isolated and monitored.

Why it mattersModel gateways, coding agents and vector databases often hold API keys, source code and private documents. Hacks of them shown at a top contest mean they should be secured as carefully as any other production system.

👀 What to watch

Watch for LiteLLM, Chroma and Oracle security advisories and patched releases, any statement from OpenAI on Codex, and ZDI's final results, including whether Nvidia Dynamo and Chroma were successfully exploited.

3 opportunities from this story

1

AI infrastructure patch-watch and hardening service

Medium⏱ 2-3 weeks💰 Monthly retainer priced by number of components, plus a one-off hardening fee

Offer a monthly service that lists a client's self-hosted AI components, such as LiteLLM, Chroma, inference servers and agent runners, checks how they are configured, and tracks advisories so patches go in within days rather than months. With 90-day disclosure clocks now running, many startups and mid-sized companies running these tools have no one assigned to keep them updated. The buyers are CTOs and heads of platform at companies that adopted AI quickly.

Best for
DevOps and security freelancers, managed-service providers
First step this week
Write a hardening checklist for LiteLLM and Chroma (network exposure, authentication, key storage, version pinning, logging) and offer a free 30-minute review to five companies that use them.
Open the full playbook
Launch steps
  1. Build an inventory script that finds AI components and their versions
  2. Write hardening checklists for the most common tools
  3. Set up advisory monitoring from GitHub security advisories and vendor feeds
  4. Run pilots with two or three clients and record time-to-patch
  5. Turn the service into tiered plans with response-time guarantees
Tools
GitHub Security AdvisoriesTrivy or GrypeDocker/KubernetesUptime and log monitoring (e.g. Grafana)
Risks

You could be blamed if a client is breached. Limit your liability by contract, document every recommendation and the date it was made, and don't promise that systems are secure.

2

Open-source argument-injection tester for coding agents

Medium⏱ 4-6 weeks💰 Open-core: free command-line tool plus a paid CI and dashboard tier, and paid audits

Build a test suite that sends malicious inputs, such as crafted filenames, branch names and command flags, to a coding agent's tools to check whether attacker-controlled text can become command options. The Codex exploit needed only one argument-injection bug, so teams building internal agents or developer tools have the same weakness. Release a free open-source core, then sell CI integration and reports to companies.

Best for
Security engineers and developers who build agent tooling
First step this week
Create 20 argument-injection test cases for common shell and git tools and run them against an open-source coding agent in a sandbox.
Open the full playbook
Launch steps
  1. Collect known argument-injection patterns for git, curl, tar and package managers
  2. Build a sandboxed harness that safely drives an agent's tools
  3. Publish it on GitHub with clear documentation and responsible-disclosure guidance
  4. Add a GitHub Action for CI
  5. Sell audits to companies shipping agent features
Tools
Python or GoDocker sandboxesGitHub ActionsOpen-source coding agents for testing
Risks

Big vendors may build similar checks into their own products, and offensive tooling can be misused. Focus on defensive CI use, run tests only in sandboxes, and follow responsible disclosure for anything you find in third-party tools.

3

Bug bounty research focused on AI infrastructure

High⏱ 1-3 months💰 Bounties and contest prizes, plus consulting that follows from published findings

For skilled security researchers, AI infrastructure is now a recognized bounty category. Pwn2Own paid $40,000 each for the LiteLLM, Codex and Oracle exploits. Specializing in gateways, vector stores and agent runtimes lets you build expertise in a field with less competition than browsers or phones. Income comes from contest prizes, vendor bounty programs and the consulting work that a public record often leads to.

Best for
Experienced penetration testers and vulnerability researchers
First step this week
Set up a local lab with current LiteLLM and Chroma releases and review their authentication, input handling and code-execution paths, following each project's security policy for reporting.
Open the full playbook
Launch steps
  1. Read each target's security policy and how to report safely
  2. Map the attack surface: admin APIs, plugins, file handling, tool execution
  3. Report findings through official channels and track CVEs
  4. Publish write-ups after disclosure to build a reputation
  5. Use that reputation to win audit contracts
Tools
Burp SuiteSemgrep or CodeQLDocker lab environmentsVendor bug bounty platforms
Risks

Bounty income is irregular, and duplicate reports (like the LiteLLM 'collision') earn less. Treat bounties as a way to build a portfolio alongside steady consulting, and test only within authorized scope.

Get these as a PDF every 3 days

Free. One email every 3 days. Unsubscribe any time.

More editions