停止像巨量问题解决者一样使用大语言模型
TL;DR · AI 摘要
文章主要讨论了关于使用大语言模型(LLMs)作为巨大问题解决者的常见误解和建议停止这样做。
核心要点
- 大语言模型并不擅长解决复杂问题。
- 应将大语言模型用于辅助而非替代人类工作。
- 关注模型的局限性和正确使用场景。
结构提纲
按章节快速跳转。
思维导图
用一张图看清主题之间的关系。
查看大纲文本(无障碍 / 无 JS 友好)
- 停止使用大语言模型如巨量问题解决者
- 大语言模型的局限性
- 不擅长复杂问题
- 更适合简单任务
- 正确的使用方式
- 辅助人类工作
- 了解适用范围
- 结论
- 不应过度依赖
金句 / Highlights
值得收藏与分享的关键句。
大语言模型并不擅长解决复杂问题,而是更适合处理简单的信息检索和文本生成任务。
建议将大语言模型用于辅助人类工作,而不是完全替代人类。
关注模型的局限性,了解其适用范围,以避免错误的应用。
Stop Using LLMs Like Giant Problem Solvers | Towards Data Science
We value your privacy
We use cookies to enhance your browsing experience, serve personalised ads or content, and analyse our traffic. By clicking "Accept All", you consent to our use of cookies.
Customise Reject All Accept All
Customise Consent Preferences
We use cookies to help you navigate efficiently and perform certain functions. You will find detailed information about all cookies under each consent category below.
The cookies that are categorised as "Necessary" are stored on your browser as they are essential for enabling the basic functionalities of the site. ...Show more
Necessary Always Active
Necessary cookies are required to enable the basic features of this site, such as providing secure log-in or adjusting your consent preferences. These cookies do not store any personally identifiable data.
- Cookie BCTempID
- Duration 10 minutes
- Description No description available.
- Cookie __cf_bm
- Duration 1 hour
- Description This cookie, set by Cloudflare, is used to support Cloudflare Bot Management.
- Cookie AWSALBCORS
- Duration 7 days
- Description Amazon Web Services set this cookie for load balancing.
- Cookie _cfuvid
- Duration session
- Description Cloudflare sets this cookie to track users across sessions to optimize user experience by maintaining session consistency and providing personalized services
- Cookie li_gc
- Duration 6 months
- Description Linkedin set this cookie for storing visitor's consent regarding using cookies for non-essential purposes.
- Cookie __hssrc
- Duration session
- Description This cookie is set by Hubspot whenever it changes the session cookie. The __hssrc cookie set to 1 indicates that the user has restarted the browser, and if the cookie does not exist, it is assumed to be a new session.
- Cookie __hssc
- Duration 1 hour
- Description HubSpot sets this cookie to keep track of sessions and to determine if HubSpot should increment the session number and timestamps in the __hstc cookie.
- Cookie wpEmojiSettingsSupports
- Duration session
- Description WordPress sets this cookie when a user interacts with emojis on a WordPress site. It helps determine if the user's browser can display emojis properly.
- Cookie BCSessionID
- Duration 1 year 1 month 4 days
- Description Blueconic sets this cookie as a unique identifier for the BlueConic profile.
- Cookie _octo
- Duration 1 year
- Description No description available.
- Cookie logged_in
- Duration 1 year
- Description No description available.
- Cookie __Secure-YEC
- Duration past
- Description YouTube sets this cookie to stores the user's video player preferences using embedded YouTube video
- Cookie __eoi
- Duration 6 months
- Description Description is currently not available.
- Cookie AWSALBTGCORS
- Duration 7 days
- Description No description available.
- Cookie login-status-p
- Duration past
- Description Description is currently not available.
- Cookie AWSALBTG
- Duration 7 days
- Description No description available.
- Cookie csrf_token
- Duration session
- Description No description available.
- Cookie token_v2
- Duration 1 day
- Description Description is currently not available.
- Cookie D
- Duration 1 year
- Description Description is currently not available.
- Cookie PHPSESSID
- Duration session
- Description This cookie is native to PHP applications. The cookie stores and identifies a user's unique session ID to manage user sessions on the website. The cookie is a session cookie and will be deleted when all the browser windows are closed.
- Cookie VISITOR_PRIVACY_METADATA
- Duration 6 months
- Description YouTube sets this cookie to store the user's cookie consent state for the current domain.
- Cookie cookietest
- Duration session
- Description The cookietest cookie is typically used to determine whether the user's browser accepts cookies, essential for website functionality and user experience.
- Cookie __Host-airtable-session
- Duration 1 year
- Description This cookie is used to enable us to integrate the services of Airtable.
- Cookie __Host-airtable-session.sig
- Duration 1 year
- Description This cookie is used to enable us to integrate the services of Airtable.
- Cookie m
- Duration 1 year 1 month 4 days
- Description Stripe sets this cookie for fraud prevention purposes. It identifies the device used to access the website, allowing the website to be formatted accordingly.
- Cookie BIGipServer*
- Duration session
- Description Marketo sets this cookie to collect information about the user's online activity and build a profile about their interests to provide advertisements relevant to the user.
- Cookie __cfruid
- Duration session
- Description Cloudflare sets this cookie to identify trusted web traffic.
- Cookie _GRECAPTCHA
- Duration 6 months
- Description Google Recaptcha service sets this cookie to identify bots to protect the website against malicious spam attacks.
- Cookie __Secure-YNID
- Duration 6 months
- Description Google cookie used to protect user security and prevent fraud, especially during the login process.
- Cookie cookieyes-consent
- Duration 1 year
- Description CookieYes sets this cookie to remember users' consent preferences so that their preferences are respected on subsequent visits to this site. It does not collect or store any personal information about the site visitors.
Functional
- [x]
Functional cookies help perform certain functionalities like sharing the content of the website on social media platforms, collecting feedback, and other third-party features.
- Cookie lidc
- Duration 1 day
- Description LinkedIn sets the lidc cookie to facilitate data center selection.
- Cookie brw
- Duration 1 year
- Description No description available.
- Cookie brwConsent
- Duration 5 minutes
- Description Description is currently not available.
- Cookie WMF-Uniq
- Duration 1 year
- Description Description is currently not available.
- Cookie loom_anon_comment
- Duration 1 year
- Description No description available.
- Cookie loom_referral_video
- Duration session
- Description Description is currently not available.
- Cookie VISITOR_INFO1_LIVE
- Duration 6 months
- Description A cookie set by YouTube to measure bandwidth that determines whether the user gets the new or old player interface.
- Cookie yt-remote-connected-devices
- Duration Never Expires
- Description YouTube sets this cookie to store the user's video preferences using embedded YouTube videos.
- Cookie ytidb::LAST_RESULT_ENTRY_KEY
- Duration Never Expires
- Description The cookie ytidb::LAST_RESULT_ENTRY_KEY is used by YouTube to store the last search result entry that was clicked by the user. This information is used to improve the user experience by providing more relevant search results in the future.
- Cookie yt-remote-device-id
- Duration Never Expires
- Description YouTube sets this cookie to store the user's video preferences using embedded YouTube videos.
- Cookie yt-remote-session-name
- Duration session
- Description The yt-remote-session-name cookie is used by YouTube to store the user's video player preferences using embedded YouTube video.
- Cookie yt-remote-fast-check-period
- Duration session
- Description The yt-remote-fast-check-period cookie is used by YouTube to store the user's video player preferences for embedded YouTube videos.
- Cookie yt-remote-session-app
- Duration session
- Description The yt-remote-session-app cookie is used by YouTube to store user preferences and information about the interface of the embedded YouTube video player.
- Cookie yt-remote-cast-available
- Duration session
- Description The yt-remote-cast-available cookie is used to store the user's preferences regarding whether casting is available on their YouTube video player.
- Cookie yt-remote-cast-installed
- Duration session
- Description The yt-remote-cast-installed cookie is used to store the user's video player preferences using embedded YouTube video.
- Cookie cp_session
- Duration 3 months
- Description Codepen sets this cookie for Help systems found in the website.
- Cookie loid
- Duration 1 year 1 month 4 days
- Description This cookie is set by the Reddit. The cookie enables the sharing of content from the website onto the social media platform.
Analytics
- [x]
Analytical cookies are used to understand how visitors interact with the website. These cookies help provide information on metrics such as the number of visitors, bounce rate, traffic source, etc.
- Cookie __hstc
- Duration 6 months
- Description Hubspot set this main cookie for tracking visitors. It contains the domain, initial timestamp (first visit), last timestamp (last visit), current timestamp (this visit), and session number (increments for each subsequent session).
- Cookie hubspotutk
- Duration 6 months
- Description HubSpot sets this cookie to keep track of the visitors to the website. This cookie is passed to HubSpot on form submission and used when deduplicating contacts.
- Cookie _ga
- Duration 1 year 1 month 4 days
- Description Google Analytics sets this cookie to calculate visitor, session and campaign data and track site usage for the site's analytics report. The cookie stores information anonymously and assigns a randomly generated number to recognise unique visitors.
- Cookie _ga_*
- Duration 1 year 1 month 4 days
- Description Google Analytics sets this cookie to store and count page views.
- Cookie __Host-psifi.analyticsTrace
- Duration 6 hours
- Description Description is currently not available.
- Cookie __Host-psifi.analyticsTraceV2
- Duration 6 hours
- Description Description is currently not available.
- Cookie _gh_sess
- Duration session
- Description GitHub sets this cookie for temporary application and framework state between pages like what step the user is on in a multiple step form.
- Cookie YSC
- Duration session
- Description YSC cookie is set by Youtube and is used to track the views of embedded videos on Youtube pages.
- Cookie ajs_anonymous_id
- Duration 1 year
- Description This cookie is set by Segment to count the number of people who visit a certain site by tracking if they have visited before.
- Cookie vuid
- Duration 1 year 1 month 4 days
- Description Vimeo installs this cookie to collect tracking information by setting a unique ID to embed videos on the website.
Performance
- [x]
Performance cookies are used to understand and analyse the key performance indexes of the website which helps in delivering a better user experience for the visitors.
- Cookie AWSALB
- Duration 7 days
- Description AWSALB is an application load balancer cookie set by Amazon Web Services to map the session to the target.
- Cookie acq
- Duration past
- Description Description is currently not available.
- Cookie acq.sig
- Duration past
- Description Description is currently not available.
- Cookie ptc
- Duration 2 years
- Description No description available.
Advertisement
- [x]
Advertisement cookies are used to provide visitors with customised advertisements based on the pages you visited previously and to analyse the effectiveness of the ad campaigns.
- Cookie muc_ads
- Duration 1 year 1 month 4 days
- Description Twitter sets this cookie to collect user behaviour and interaction data to optimize the website.
- Cookie guest_id_marketing
- Duration 1 year 1 month 4 days
- Description Twitter sets this cookie to identify and track the website visitor.
- Cookie guest_id_ads
- Duration 1 year 1 month 4 days
- Description Twitter sets this cookie to identify and track the website visitor.
- Cookie personalization_id
- Duration 1 year 1 month 4 days
- Description Twitter sets this cookie to integrate and share features for social media and also store information about how the user uses the website, for tracking and targeting.
- Cookie guest_id
- Duration 1 year 1 month 4 days
- Description Twitter sets this cookie to identify and track the website visitor. It registers if a user is signed in to the Twitter platform and collects information about ad preferences.
- Cookie bcookie
- Duration 1 year
- Description LinkedIn sets this cookie from LinkedIn share buttons and ad tags to recognize browser IDs.
- Cookie __Secure-ROLLOUT_TOKEN
- Duration 6 months
- Description YouTube sets this cookie to manage feature rollout and experimentation. It helps Google control which new features or interface changes are shown to users as part of testing and staged rollouts, ensuring consistent experience for a given user during an experiment.
- Cookie yt.innertube::nextId
- Duration Never Expires
- Description YouTube sets this cookie to register a unique ID to store data on what videos from YouTube the user has seen.
- Cookie yt.innertube::requests
- Duration Never Expires
- Description YouTube sets this cookie to register a unique ID to store data on what videos from YouTube the user has seen.
- Cookie session_tracker
- Duration session
- Description This cookie is set by the Reddit. This cookie is used to identify trusted web traffic. It also helps in adverstising on the website.
- Cookie edgebucket
- Duration session
- Description Reddit sets this cookie to save the information about a log-on Reddit user, for the purpose of advertisement recommendations and updating the content.
- Cookie did
- Duration 1 year
- Description Arbor sets this cookie to show targeted ads to site visitors.This cookie expires after 2 months or 1 year.
Uncategorised
Other uncategorised cookies are those that are being analysed and have not been classified into a category as yet.
No cookies to display.
Reject All Save My Preferences Accept All
Publish AI, ML & data-science insights to a global community of data professionals.
- * *
Toggle Mobile Navigation
Toggle Search
Search
Stop Using LLMs Like Giant Problem Solvers
How I turned 100 messy pdfs into structured insights by building a deterministic loop around agents
May 26, 2026
6 min read
Share

Image by Wesley Tingey from Unsplash
I recently worked on a feature where I had to transform 100 messy compliance pdfs into structured JSON rules.
The brute force approach was obvious: give the agent the source text, explain the task, provide examples, and ask it to generate the rules. Since it was the lowest-hanging fruit, I tried it first.
At a glance, the output looked fine. The output JSON was valid and matched what I expected.
But as I was manually sampling the results to check for accuracy, the cracks appeared. Some rules were too broad, others were missed. Some rules failed to preserve the nuances of the original text. I tried using another agent to catch and fix the errors but with such a huge corpus, it was impossible to confidently verify the output.
That was the frustrating part. The errors were not obvious. This was way too fragile of an implementation to scale.
Though I cannot share the exact implementation details, what I can share are the architectural lessons I learnt and how I eventually implemented it. Hopefully, these insights will be useful if you’re building AI systems that need to scale, stay reliable, and deal with messy data. And if you have better ways of doing things, do reach out to chat!
Okay let’s get to it.
The problem
The 100 pdfs I worked with had already been parsed and chunked before they reached me. But the raw content was still messy. There were bullet points, tables, OCR artefacts, translated sections, semi structured headings, footers, headers, inconsistent formatting and document specific quirks.
I chose to use an agent because deciding what mattered required semantic judgement. The documents did not follow one consistent pattern, so relevance could not be determined through simple rules alone.
You had to understand the surrounding context. None of this was difficult when done on a small chunk of data. The challenge was performing this reliably at scale.
These rules were then processed by another downstream system to be evaluated deterministically.
What eventually worked
After a few experiments, I realised the biggest improvement did not come from a better prompt, a new tool, an MCP server, or a more sophisticated agent harness.
It came from changing the shape of the problem.
Instead of trying to make the agent smarter, I made the agent’s job smaller.
The first change was to prepare the source data upfront. Instead of asking the agent to query a database, retrieve records, decide whether it had the right inputs, and then perform the extraction, I gave it a more controlled starting point.
In my case, that meant temporarily storing the relevant raw data locally.
This may not always be practical. But the underlying principle is to reduce the amount of retrieval uncertainty the agent has to handle. If the agent’s job is to reason over content, do not also make it responsible for figuring out whether it has found the right content.
Another option would be to prepare the query upfront.
I also used a script to strip away unnecessary metadata and fields before passing the raw content to the agent. Less irrelevant context meant fewer distractions, fewer chances for the agent to latch onto the wrong details and a cleaner reasoning task overall.
But the most important change was the unit of work.
Instead of processing everything at once, I did things iteratively and processed one document at a time.
That made each job smaller, easier to inspect, easier to retry, and easier to audit. I spun up five subagents to process documents in parallel, with each agent logging its progress to a file.
If one document failed, I could retry only that document. If one output had formatting issues, I could fix that specific case without rerunning the whole batch. If the pipeline stopped halfway, the cached progress meant it could resume from the last successful checkpoint.
This was also where the separation of responsibilities became clearer.
The agent handled the semantic work: understanding the content, identifying the relevant parts and writing the JSON output.
The surrounding code handled the mechanical parts: parallelising jobs, enforcing the schema, generating IDs, writing files, caching progress, validating references, and checking whether the output could be traced back to the original source.
I also had an orchestrator watch over the progress of the script.
Making the output auditable
A useful design decision was adding reference IDs to every generated rule. This meant that each output item pointed back to a specific source.
This made the output easier to audit. Instead of asking, “Does this generated rule look right?”, I could ask more precise questions such as: does the referenced source chunk exist? Is the quoted source text actually present in that chunk?
I could also get another agent to selectively run audits on larger and more complex documents to ensure that important nuances were preserved.
On top of that, I did a lightweight version of evals. I ran a small batch of raw documents through the workflow and manually reviewed the results for coverage and accuracy. A full golden dataset was not practical for the scope of this task, but I still needed a way to prove to myself that the workflow was working.
My goal was not to build a perfect benchmark but to make the system auditable enough that I could inspect the outputs, catch failures, and iterate toward a higher accuracy bar.
If you’ve got ideas on how I could have done this better, let me know!
My biggest takeaway
The pattern that worked was to stop treating the LLM as the whole system.
The system became more reliable not because the agent became perfect, but because the workflow made its outputs easier to trace, validate, and recover from.
Coincidentally, I was building this shortly before attending the inaugural AI Engineer Singapore conference, held from 15–17 May 2026.
On the last day, JJ Geewax, Director of Applied AI at Google DeepMind, shared a framing that captured what I had been learning the hard way: we need to stop using LLMs like giant problem solvers.
That resonated with me because it is such an easy trap to fall into. It is easy to just give the model the data, schema, business rules, edge cases, and the responsibility to verify itself. Then get frustrated when the result is inconsistent.
But for reliable production systems, the better pattern is usually a hybrid. Let the agent handle the parts that require semantic judgement, and let code handle the parts that require structure, validation, and control.
I’ll be sharing more reflections from AI Engineer Singapore and the workshops I attended. The YouTube snippet of JJ’s speech here.
That’s all from me. I hope this helped, and see you in the next article
- * *
Written By
Clara Chong
Agentic Ai, Artificial Intelligence, Editors Pick, Llm, Programming
Share This Article
Towards Data Science is a community publication. Submit your insights to reach our global audience and earn through the TDS Author Payment Program.
Related Articles
Natural Language Processing How machines make sense of sentence structure: Combinatory Categorial Grammar Marco Hening Tallarico June 2, 2025 12 min read
Artificial Intelligence The exciting new world of designing conversation driven APIs for LLMs. Roni Dover July 28, 2025 9 min read
Artificial Intelligence A framework for measuring retrieval quality in Model Context Protocol agents. Tomaz Bratanic July 29, 2025 9 min read
LLM Applications From random example selection to systematic AuPair generation — how to make your LLM prompts actually… Sudheer Singh August 7, 2025 7 min read
LLM Applications A primer on overcoming LLM limitations with formal verification. Jacopo Tagliabue August 20, 2025 12 min read
LLM Applications Explore how to transcribe videos with speaker identification in a single prompt Laurent Picard August 29, 2025 66 min read
LLM Applications Exploring ways to make voice assistants more personal Deepak Krishnamurthy August 31, 2025 8 min read
Your home for data science and Al. The world’s leading publication for data science, data analytics, data engineering, machine learning, and artificial intelligence professionals.
© Insight Media Group, LLC 2026
Subscribe to Our Newsletter
Some areas of this page may shift around if you resize the browser window. Be sure to check heading and document order.
##
##
##
##
##
##
##