Non-Dummy variable codings in `ContinuousEncoder`

Right now, MLJModels.jl has ContinuousEncoder, which automatically transforms into one of two codings:
1. Dummy variable coding (one-hot with last category dropped)
2. Redundant variable coding (one-hot)

However, these aren't always the most useful codings for effective regularization, and [there are many others in common use](https://juliastats.org/StatsModels.jl/stable/contrasts/). For example, the most common way to encode ordinal variables is with sequential difference encoding; with this encoding, regularization pulls adjacent categories closer together, which improves model performance relative to either treating ordered variables as categorical (discarding ordering information) or treating them as continuous (using an equal-distance assumption that is often incorrect). Similarly, effect coding allows you to regularize categories towards the grand mean (rather than regularize every category towards 0, or regularize all categories towards one other category).

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Non-Dummy variable codings in `ContinuousEncoder` #534

Metadata

Assignees

Labels

Type

Projects

Milestone

Relationships

Development

Non-Dummy variable codings in ContinuousEncoder #534

Description

Metadata

Metadata

Assignees

Labels

Type

Projects

Milestone

Relationships

Development

Issue actions

Non-Dummy variable codings in `ContinuousEncoder` #534