Talks
- From Agent Traces to Agent Simulations — Rustem Feyzkhanov, Snorkel AI
- Setting Yourself Up for Success with Codex — Jason Liu (OpenAI Workshop)
- From Signal to PR: Anatomy of a Self-Improving Agent — Jason Lopatecki, Arize
- Claude for Long-Horizon Tasks — Lance Martin, Anthropic
- Why Agentic Systems Need Ontologies — Frank Coyle, UC Berkeley
- Don't Ship Skills Without Evals — Philipp Schmid, Google DeepMind
- Eve Bouffard's AI-First Design Workflow: Paxel, SOTA Zine, and Startup School
- The CEO Must Be the Chief AI Officer — Pedro Franceschi (Y Combinator Lightcone)
- LLM Observability, Evaluation, and Experimentation for AI Agents — Dat Ngo, Arize AI
- How Anthropic Uses Claude in GTM Engineering: Building CLAFTS with Claude Code
- How to Build a Self-Improving Company with AI (Tom Blomfield, YC)
- Inside YC's AI Playbook: Pete Koomen on Building Agent Infrastructure at Y Combinator
- Harness Engineering: How to Build Software When Humans Steer, Agents Execute — Ryan Lopopolo, OpenAI
- Reflecting on a Year of Claude Code — Boris Cherny & Cat Wu (2025)
- To Give an Agent a Workspace: Laying a Foundation for Messy Forever Work
- RAG Is Dead, Right? — Kuba Rogut, Turbopuffer
- Build a Proactive Agent Workflow with Claude Code (Routines)
Everything below is generated from the talks — the nodes of the knowledge graph, each page accumulating what speakers say across talks.
Contradictions
Concepts
- Ablation Testing
- Agent Infrastructure
- Agent Ops
- Agent Simulation
- Agent Skills
- Agent Workspace
- Agent-on-Agent Review
- Agentic Loops
- Agentic Search
- AI as Feature vs AI as OS
- AI Email Assistant
- AI loop architecture
- Alex Decision Lens
- Append-only event log
- Ask-Plan-Audit-Build-Verify Workflow
- Atomic PRs
- Auto Mode
- Auto-compaction
- Business metrics
- Capability Skills
- Centralization vs Decentralization of AI
- Chat Interface
- CI Pipeline for Agents
- CLAUDE.md
- Codebase Uniformity
- Compaction
- Company AGI
- Company brain
- Computer Use
- Context Minimalism
- Context Retrieval
- Context window
- Continual Learning Bench
- Copilot model
- Corporate AI
- Customer World Model
- Decoupling brain and hands
- Denormalization for Agents
- Deploy Verifier
- Deterministic evals
- Diarization
- Discussion Gate
- disposable design
- Domain knowledge extraction
- Dreaming
- DRY
- Electricity Analogy
- Embeddings
- Evals
- Experimentation
- Feature Ledger
- Found Work
- Full Text Search
- Garbage collection day
- Gen 4 Cloud Workspace
- Golden datasets
- GTM Engineering
- Hallucination
- Harbor Format
- Harness Engineering
- Hive Mode
- Horseless Carriages
- HTTP Proxy Security
- Human feedback
- human vs machine web design
- Human-in-the-Loop
- imagination as bottleneck
- In-band memory writing
- Individual contributor
- Intel Corpus
- Iterative Retrieval
- Jevons Paradox
- Journal Entries
- Just-in-Time Software
- Knowledge Graph
- Launch Contract
- LLM
- LLM as a judge
- LLM as fuzzy compiler
- LLM Wikis
- locally personalized software
- Long-horizon tasks
- Loop
- Managed Agents
- MCP
- MECE
- Memory substrate
- Memory Vault
- Merkle Trees
- Messy Forever Work
- Middle management
- Minimal Surface Area
- Model-Triggered Skills
- Monitoring agent
- Monorepo
- mood board
- Multi-span evals
- Multiplayer Harness
- Neurosymbolic AI
- No-ops
- Non-functional requirements
- Observability
- Observability Traces
- one-shot prototyping
- Online Evals
- Ontology
- Ontology Validator
- Operational AI
- Oracle Solution
- Org-level harness
- paper shaders
- Parallel Work Streams
- Parameter golf
- Personal AI Revolution
- Pinned Threads
- Plan Mode
- Plan of Plans
- Post-training
- Preference Skills
- Private Benchmark
- Proactive Agent
- Product AI
- Progressive disclosure
- Prompt injection
- Proof of Concept (PC)
- Proof Receipts
- Public Benchmark
- RAG
- Red Teaming
- Resolution
- Resolver
- Review agents
- Reward Hacking
- Roadmap Folder
- Role Merging
- Routines
- Sandboxes
- Self-Improving Agents
- Self-improving company
- Self-Improving Dream Cycle
- Self-Improving Systems
- send-to-agent feature request
- Session-level evals
- shader fine-tuning modal
- Shared Organizational Brain
- Ship Flow
- Simulated User
- Single-Player Era of Agents
- Skill Lab
- Skill Retirement
- Skillify
- Skills
- Software ephemerality
- soul.md
- Span evals
- Spec
- SQL Agent Access
- Sub-agents
- Subject Matter Expert
- Symbolic AI
- Task horizon
- Ticket
- Token usage
- Tool Registry
- Tool Use Loop
- Traces and spans
- Trajectory evals
- Trust-Default Culture
- Two-Sentence Pitch Skill
- User-Invoked Skills
- Vector Search
- Verification
- Verifier
- Verifier loop
- Virtual Employee
- VPC deployment
- WWK (What We Know)
- YC User Manual
People
- Aaron Epstein
- Alex Damis
- Andrej Karpathy
- Boris Cherney
- Boris Cherny
- Cat Wu
- Dat Ngo
- Eve Bouffard
- Frank Coyle
- Garry Tan
- Harj
- Jack Dorsey
- Jared
- Jared Sires
- Jason Liu
- Jason Lopatecki
- Jeff Dean
- Kuba Rogut
- Lance Martin
- Maya
- Nate B. Jones
- Pedro Franceschi
- Pete Koomen
- Philipp Schmid
- Rustem Feyzkhanov
- Ryan Lopopolo
- Tom Blomfield
- Vibhu Sapra
Tools
- Agent SDK
- Agent View
- Alyx
- Appshots
- Aqua
- Arize Phoenix
- CD Pipeline
- CLAFTS
- Claude
- Claude API
- Claude Code
- Claude Managed Agents
- Claude Opus 4
- Claude Tag
- Conductor
- Context7
- Crab Trap
- Cron
- Cursor
- DataDog
- Daytona
- ESLint
- ffmpeg
- Figma
- GBrain
- Gemini CLI
- Gemini Interactions API
- GitHub
- Google Docs
- Google Drive
- Grafana
- Hermes Agent
- Magpie
- Messages API
- OpenClaw
- OpenTelemetry
- Optimizely
- OWL
- Paper Design
- Paxel
- Pi
- Postgres
- Pydantic
- Pyroscope
- RDFS
- Remote Control
- Schema.org
- Signal
- Skills Bench
- Slack
- SOTA Zine
- Super Whisper
- SWE-bench
- TerminalBench
- Twilio
- Voice Mode
- WebArena
- Windsurf
- YC Startup School
- Zod