Skip to content

fix(eval): strip query strings when extracting Wikipedia titles from URLs - #168

Open
longzhenren wants to merge 1 commit into
StarTrail-org:mainfrom
longzhenren:fix/wiki-title-strip-query
Open

longzhenren wants to merge 1 commit into
StarTrail-org:mainfrom
longzhenren:fix/wiki-title-strip-query

Conversation

@longzhenren

Copy link
Copy Markdown

Root cause

WikipediaAPIRetriever._extract_wiki_title strips only the #fragment suffix from Wikipedia URLs:

pattern = r"https?://[a-z]{2,3}\.wikipedia\.org/wiki/(.+?)(?:#.*)?$"

Mobile share links almost always carry query parameters (e.g. ?wprov=rarw1), so the extracted title kept the query string and the downstream retrieve() call looked up a non-existent article, degrading retrieval quality.

Repro on main:

url = "https://en.wikipedia.org/wiki/Albert_Einstein?wprov=rarw1"
# before: "Albert Einstein?wprov=rarw1"
# after:  "Albert Einstein"

Fix

Strip both query and fragment suffixes: the pattern tail becomes (?:[?#].*)?$.

Test

Added tests/test_wiki_title_extraction.py covering: plain URLs, query-string stripping, fragment stripping, percent-encoded titles, and non-Wikipedia URLs returning None.

…URLs

Mobile share links commonly carry query parameters like ?wprov=, which
ended up inside the extracted title and made retrieve() look up a
non-existent article. Strip both ?query and #fragment suffixes.
@vercel

vercel Bot commented Oct 2, 2026

Copy link
Copy Markdown

@longzhenren is attempting to deploy a commit to the andylizf's projects Team on Vercel.

A member of the Team first needs to authorize it.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant