Sort Characters By Frequency
Rearrange a string so characters appear in blocks ordered by descending frequency.
- 1 <= s.length <= 5 * 10⁵
- s consists of uppercase and lowercase English letters and digits.
Intuition
Sort characters by frequency rearranges a string so characters appear in descending order of how often they occur, with all copies of each character grouped together.
The shape is direct — count, then order by count:
- Tally each character's frequency, sort the distinct characters by that frequency descending, then emit each character repeated its count times.
Sorting the distinct characters rather than the string itself is what keeps this efficient. There are at most k distinct characters, so the sort is O(k log k) rather than O(n log n) — a real difference when the string is long and its alphabet small.
The output must group identical characters, not interleave them. Emitting each character's full run at once satisfies that automatically.
Characters with equal frequency may appear in any relative order, so no tie-breaking rule is needed — a detail worth confirming against the problem statement rather than assuming.
A bucket sort avoids comparison sorting entirely. Frequencies range from 1 to n, so bucketing characters by count and reading the buckets from high to low is O(n). That is asymptotically better, though the comparison sort is simpler and fast enough given the small alphabet.
Building the result with repeated string concatenation is O(n²) in languages with immutable strings. Appending to a list and joining once keeps it linear.
The input may include uppercase, lowercase, and digits, so a hash map is safer than a fixed 26-element array here.
Counting is O(n), sorting O(k log k), and building O(n), giving O(n + k log k) overall.
Count, sort the distinct characters by count descending, then emit each one repeated. Sorting the 26-ish distinct keys rather than the string's characters is what keeps this near-linear in the input length.
Approach
Before reading on: price up what counting everything costs here, then ask why only the extreme element matters and the rest need no ordering. Aim for O(n + k log k) time and O(n) space.
Count each character
Tally frequencies in one pass. A hash map is safer than a fixed array here, since the input includes uppercase, lowercase, and digits.
Sort the distinct characters
Sort the k distinct characters, not the n positions. That makes the sort O(k log k) rather than O(n log n) on a long string.
Emit each run together
Write every copy of a character consecutively. The output must group identical characters, which emitting full runs satisfies automatically.
Ignore ties
Characters with equal frequency may appear in any relative order, so no tie-breaking is required — worth confirming rather than assuming.
Consider bucket sort
Frequencies range from 1 to n, so bucketing by count and reading high to low is O(n), avoiding comparison sorting altogether.
Build with a join
Repeated concatenation is O(n²) on immutable strings. Append to a list and join once to keep the construction linear.
Cost of the approach
Counting is O(n), sorting O(k log k), building O(n) — O(n + k log k) overall with O(n) space.
Solution & live demo
Common pitfalls
Sorting the characters of the string
return ''.join(sorted(s, key=lambda c: -counts[c]))
ordered = sorted(counts, key=lambda c: -counts[c])
That's O(n log n) on the full string rather than O(k log k) on the distinct keys. It also relies on sort stability to keep equal-frequency characters grouped, which is fragile reasoning.
Building the result character by character
for c in ordered:
for _ in range(counts[c]): result.append(c)result.append(c * counts[c])
String repetition emits the whole run in one operation. The inner loop does the same work with per-character overhead and more code.
Sorting ascending
sorted(counts, key=lambda c: counts[c])
sorted(counts, key=lambda c: -counts[c])
The problem asks for decreasing frequency, so the most common character must come first. The negation (or reverse=True) is the whole ordering requirement.
Edge cases
the frequency map is empty, the loop over distinct characters does nothing, and the result is the empty string
every count is 1, so any order among them is a valid answer since ties break arbitrarily
a single distinct entry with a large count still emits correctly as one block
'A' and 'a' are different keys in the frequency map, since character equality is case-sensitive