Identifying anomalies or outliers in a set of data records employs a distance or similarity measure between features of record pairs that depends upon the frequencies of the feature values in the set. According to this instruction, the erasion level is set. This determination can be based upon any of a variety of decision-making criteria as may pertain to a given application setting.