Stop ignoring numbers without a unit in relative dates - #1410
AdrianAtZyte wants to merge 5 commits into
Conversation
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## master #1410 +/- ##
=======================================
Coverage 98.94% 98.95%
=======================================
Files 239 239
Lines 3521 3528 +7
=======================================
+ Hits 3484 3491 +7
Misses 37 37 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
serhii73
left a comment
There was a problem hiding this comment.
Thanks! This fixes #1034 (03/06/2020a gives 2020-06-03). The catch is that a relative date followed
by any unrelated number is now rejected, and in search_dates that is common prose:
from dateparser.search import search_dates
for s in ["Posted 3 days ago 15 comments", "John Doe 2 hours ago 3 replies",
"Reviewed 3 weeks ago 12 people found this helpful", "Answered 2 years ago 7 upvotes"]:
print(s, search_dates(s, languages=["en"]))On master these give 3 days / 2 hours / 3 weeks / 2 years ago. On this branch (no RELATIVE_BASE)
they give 6 days / ~13 hours / 6 weeks / 4 years ago: the offset is computed from a wrong base. Once 3 day ago 15 is
rejected, search splits the chunk. parse_item parses 3 day in one candidate split and sets
RELATIVE_BASE on parser._settings without restoring it, so the 3 day ago of the next split is
computed from that. The leak is old, but this change is what sends ordinary text through it. With a
RELATIVE_BASE, some chunks are lost instead: "Posted 3 days ago | 15 comments" and
"Last seen 3 days ago at 5" go from 3 days ago to None.
The same rejection hits auto-detected French: search_dates("Il y a 5 min") is translated as
year 1 5 minute, and the stray 1 now makes it January 5th (master: 5 minutes ago). parse() also
loses the day for yesterday at 5, ayer a las 9, вчера в 12 or 3 days ago 10.30 (#1450 recovers
some of these).
Would it work to keep rejecting bare numbers before the quantity (the #1034 shape) and ignore the ones
after the last one? For example, in _are_all_words_units:
matches = list(PATTERN.finditer(date_string))
if matches:
end = matches[-1].end()
date_string = date_string[:end] + re.sub(r"\d+", "", date_string[end:])
date_string = PATTERN.sub("", date_string)With that, the tests in this PR still pass, 03/06/2020a and the 1st of last month are still handled,
and the search_dates and parse() examples above go back to master's results. The
parser._settings leak in search.py parse_item is probably worth fixing on its own too: restoring
the settings after the re-parse also fixes the doubled offsets.
Fixes #1034. Part of #1265.