Real Python Podcast E304 Title Image

Episode 304: Configuring a Versatile LLM Harness & Scraping the Web With Scrapy

The Real Python Podcast

Which is more important, the model or the “harness” around an LLM? What are ways to assemble an efficient agentic developer workflow? This week on the show, Ayan Pahwa joins us to discuss harnessing, web scraping, and self-hosting Python applications.

Episode Sponsor:

Ayan is a developer advocate at Zyte and an experienced project builder. We discuss a recent article he wrote about creating an extension for the web scraping tool Scrapy. He also digs into his self-hosting setup for Python applications and tools.

Our discussion extends to the complexities of developing effective harnesses. Ayan shares his setup and how he navigated shifting from prompt engineering to context and loop engineering.

This episode is sponsored by HydraDB.

Topics:

  • 00:00:00 – Introduction
  • 00:02:01 – Scrapy and building an extension
  • 00:08:46 – Zyte and the web scraping API
  • 00:11:19 – Sponsor: HydraDB
  • 00:12:22 – noalgotube project
  • 00:15:49 – Homelab & self hosting projects
  • 00:22:19 – What goes into a harness?
  • 00:32:33 – Where did you start exploring LLM tools?
  • 00:36:13 – Local models & edge computing
  • 00:39:31 – ExtractPod and discussing Apple’s AI
  • 00:43:35 – Video Course Spotlight
  • 00:44:54 – Managing token use and tools
  • 00:52:37 – What are you excited about in the world of Python?
  • 00:55:10 – What do you want to learn next?
  • 00:56:38 – What is the best way to follow your work online?
  • 00:56:58 – Thanks and goodbye

Show Links: