Skip to content

textrecipes (development version)

Bug Fixes

  • Fixed a crash (segfault) in step_texthash() and step_dummy_hash() when num_terms = 0 was used; this now errors cleanly during prep(). (#314)

  • Fixed a bug in step_texthash() where baking with sparse = "yes" and signed = TRUE (the default) could error with “values value must not be equal to the default 0” when signed hash collisions cancelled out to a true zero. (#289)

  • Fixed a bug in step_dummy_hash() where keep_original_cols = TRUE failed to restore the original columns when collapse = TRUE, and where the intermediate collapsed column could leak into the output. (#304)

  • Fixed a bug in step_clean_levels() where a backwards-compatibility shim never actually restored the trained column names, causing bake() to silently skip cleaning for legacy trained objects. (#321)

  • Fixed a bug in step_clean_levels() where character columns were never given a trained cleaning dictionary, causing bake() to clean row values inconsistently (via janitor::make_clean_names() applied to the batch of values being baked) instead of using a fixed lookup learned at prep() time. (#320)

  • step_lda() produced non-deterministic and degenerate topic features at bake time: bake() re-fit topic assignments instead of projecting new documents onto the fitted model, and rebuilt the vocabulary from new_data alone with different pruning than was used at prep() time. bake() now reuses the vocabulary learned during prep() and projects each document independently, so topic weights sum to ~1 and no longer depend on what else is in the same bake() call. (#315)

  • Custom tokenizer functions passed to custom_token are now called positionally, so functions whose first argument isn’t named x no longer error with “unused argument” at bake time. (#248)

  • step_tokenize(), step_tokenize_bpe(), step_tokenize_sentencepiece(), and step_tokenize_wordpiece() now error informatively if a tokenizer returns a list whose length doesn’t match the number of rows in the input, instead of silently recycling and duplicating rows. (#318)

  • Fixed a bug where bake() could error with “unused argument” for step_tokenize(engine = "spacyr"), step_tokenize(engine = "tokenizers.bpe"), step_tokenize_bpe(), and step_tokenize_sentencepiece() when options was non-empty, because the underlying tokenizer closures didn’t accept additional arguments. (#319)

  • step_stem() no longer silently drops the lemma/pos attributes carried by a tokenlist, so it can now be safely used upstream of step_lemma() or step_pos_filter() (#327).

  • Fixed first_person(), first_personp(), second_person(), second_personp(), third_person(), to_be(), and prepositions() (used by step_textfeature()) which tested whether a word matched the entire untokenized document rather than whether the document contained that word, causing them to almost always return 0. They now tokenize the document into words before checking membership. (#316)

  • Fixed several inconsistencies in the counting functions used by step_textfeature() (#329):

    • n_uq_urls() now counts distinct full urls instead of being capped at 2 distinct matches of the literal substring “http”/“https”.
    • n_charS() and n_uq_charS() now actually exclude urls, hashtags, and mentions from the character count, as documented.
    • n_extraspaces() no longer counts a single tab or newline as “extra spaces”; it now only counts runs of 2 or more consecutive whitespace characters.
    • All counting functions now consistently return NA when given NA input.
  • Fixed show_tokens() so that the n argument’s upper bound (nrow(rec$template)) is actually validated, instead of being silently swallowed into check_number_whole()’s ... and producing a confusing, unrelated error when validation failed. Added test coverage for show_tokens(), which previously had none. (#322)

  • step_ngram() and step_sequence_onehot() now error clearly when given num_tokens, min_num_tokens, or sequence_length less than 1, and step_ngram() now errors clearly when min_num_tokens is greater than num_tokens, instead of surfacing an unrelated low-level error. (#324)

  • step_ngram() no longer mislabels non-UTF-8 (e.g. latin1) input as UTF-8, which could cause nchar(), toupper(), and other string operations on the output to fail with encoding errors. (#323)

  • Fixed bug in step_tfidf() where sublinear_tf = TRUE could produce negative TF-IDF values because the sublinear transform was applied after normalization instead of before. (#317)

textrecipes 1.1.0

CRAN release: 2025-03-18

Improvements

textrecipes 1.0.7

CRAN release: 2025-01-23

Improvements

  • Documentation for tidy methods for all steps has been improved to describe the return value more accurately. (#262)

  • Calling ?tidy.step_*() now sends you to the documentation for step_*() where the outcome is documented. (#261)

  • step_textfeatures() has been made faster and more robust. (#265)

Bug Fixes

textrecipes 1.0.6

CRAN release: 2023-11-15

  • textfeatures has been removed from Suggests. (#255)

  • step_textfeatures() no longer returns a politeness feature. (#254)

textrecipes 1.0.5

CRAN release: 2023-10-20

textrecipes 1.0.4

CRAN release: 2023-08-17

Improvements

  • step_clean_names() now throw an informative error if needed non-standard role columns are missing during bake(). (#235)

  • The keep_original_cols argument has been added to step_tokenmerge. This change should mean that every step that produces new columns has the keep_original_cols argument. (#242)

  • Many internal changes to improve consistency and slight speed increases.

Bug Fixes

textrecipes 1.0.3

CRAN release: 2023-04-14

Improvements

  • Steps with tunable arguments now have those arguments listed in the documentation.

  • All steps that add new columns will now informatively error if name collision occurs.

Bug Fixes

  • Fixed bug where step_tf() wasn’t tunable for weight argument.

textrecipes 1.0.2

CRAN release: 2022-12-21

textrecipes 1.0.1

CRAN release: 2022-10-06

  • Examples no longer include if (require(...)) code.

textrecipes 1.0.0

CRAN release: 2022-07-02

  • Indicate which steps support case weights (none), to align documentation with other packages.

textrecipes 0.5.2

CRAN release: 2022-05-04

  • Remove use of okc_text in vignette

  • Fix bug in printing of tokenlists

textrecipes 0.5.1

CRAN release: 2022-03-29

textrecipes 0.5.0

CRAN release: 2022-03-20

New steps

Improvements and Other Changes

Breaking Changes

  • Steps with prefix argument now creates names according to the pattern prefix_variablename_name/number. (#124)

textrecipes 0.4.1

CRAN release: 2021-07-11

Bug fixes

textrecipes 0.4.0

CRAN release: 2020-11-12

Breaking Changes

  • step_lda() now takes a tokenlist instead of a character variable. See readme for more detail.

New Features

textrecipes 0.3.0

CRAN release: 2020-07-08

  • stopwords package have been moved from Imports to Suggests.
  • step_ngram() gained an argument min_num_tokens to be able to return multiple n-grams together. (#90)
  • Adds step_text_normalization() to perform unicode normalization on character vectors. (#86)

textrecipes 0.2.3

CRAN release: 2020-05-22

textrecipes 0.2.2

CRAN release: 2020-05-10

  • step_word_embeddings() got a argument aggregation_default to specify value in cases where no words matches embedding.

textrecipes 0.2.1

CRAN release: 2020-05-04

textrecipes 0.2.0

CRAN release: 2020-04-14

textrecipes 0.1.0

CRAN release: 2020-03-05

textrecipes 0.0.2

CRAN release: 2019-09-07

  • Custom stemming function can now be used in step_stem using the custom_stemmer argument.
  • step_textfeatures() have been added, allows for multiple numerical features to be pulled from text.
  • step_sequence_onehot() have been added, allows for one hot encoding of sequences of fixed width.
  • step_word2vec() have been added, calculates word2vec dimensions.
  • step_tokenmerge() have been added, combines multiple list columns into one list-columns.
  • step_texthash() now correctly accepts signed argument.
  • Documentation have been improved to showcase the importance of filtering tokens before applying step_tf() and step_tfidf().

textrecipes 0.0.1

CRAN release: 2018-12-17

First CRAN version