ASMR Editing Guide
Key Takeaways
- Noise reduction above roughly 6 to 8 dB starts eating the high frequency texture that makes ASMR tingle, and the result is the watery, artificial sound creators complain about constantly.
- Compression on ASMR should stay gentle, around a 2:1 ratio with a slow attack near 20ms, because fast heavy compression flattens the transients that trigger the response.
- Manual clip gain on individual loud sounds beats any automatic tool. Smoothing out protruding sounds by hand keeps texture a limiter would crush.
- Editing on a phone in CapCut is a real workflow, not a compromise. Plenty of working ASMR creators edit that way 99% of the time.
- Export ASMR audio at 48kHz and keep peaks under -3dB, since platform loudness normalization pulls quiet ASMR up and exposes any clipping you left in.
Editing an ASMR video is mostly an audio job with a picture attached. The tingle comes from high frequency detail between roughly 2kHz and 10kHz, tiny transients from taps, brush bristles, and breath, arriving at slightly different times in each ear.
Every editing decision either protects that detail or destroys it. Most bad ASMR edits are not under-edited, they are over-processed.
Here is the honest version of the workflow. Record clean, cut for pacing, fix noise gently, balance levels by hand, then get out. The whole thing takes 2 to 5 hours for a 30 minute video once you know what you are doing, and closer to 8 hours the first few times.
What a good ASMR edit sounds like versus a bad one
You can hear the difference in about 4 seconds. A good edit has a quiet, steady room tone underneath, no pumping, and taps that still have a crisp attack.
Listen for left to right separation on headphones if it is binaural. That separation is what makes the sound feel like it is happening in the room with you.
A bad edit sounds like it is underwater. Consonants get a swirly, phasey quality, the silence between sounds goes completely dead, and quiet parts jump in volume every few seconds.
All three are noise reduction and compression pushed too far. If your edit sounds like that, you did not need better gear, you needed less processing.
If you are editing your first video, run one test: bounce a 30 second clip processed and unprocessed, listen on the cheapest earbuds you own, and pick whichever one still has texture. If the processed one sounds smoother but flatter, back everything off by half.
The editing pass, in order
- 1. Sync and lock. If you recorded audio on a separate recorder, line up the waveform spike from a clap at the start, then group the clips so they cannot drift. Sync drift on a 45 minute video is the most common thing that ruins an otherwise finished edit.
- 2. Rough cut for pacing. Cut the dead air, the mic bumps, the setup fumbling. Leave the natural pauses. ASMR pacing is slower than any other video format, so resist trimming to normal YouTube rhythm.
- 3. Room tone patch. Grab 10 seconds of silence from your recording and lay it under every cut. Without it, each edit point makes a tiny audible hole, and viewers feel it even if they cannot name it.
- 4. Clip gain by hand. Go through and pull down the loud spikes, the plosive breaths, the one tap that jumped 6dB. This is slow and boring and it is the single highest payoff step.
- 5. Gentle noise reduction, only after clip gain, and only if you actually have hiss.
- 6. EQ. Cut, do not boost, as a default.
- 7. Light compression plus a safety limiter that almost never engages.
- 8. Color and export. Warm, dim, low contrast. Bright clinical footage fights the audio.
Noise reduction: where creators go wrong
Every noise reduction tool asks for a noise profile and a reduction amount. Take the profile from real silence in your own recording, never a preset. Then set the reduction far lower than feels right. Six dB fixes most hiss, and twelve dB is where the artifacts start.
The failure mode is specific. Over-application of noise reduction creates a watery, artificial sound, because the algorithm cannot tell hiss from the faint high frequency tail of a brush stroke. It removes both. You end up with clean silence and dead triggers.
If your room is genuinely noisy, fix the room instead. Record at night, kill the HVAC, throw a duvet over the recording space. Ten minutes of that beats an hour of repair plugins every time.
EQ and compression settings that actually suit ASMR
Start with a high pass filter at 60 to 80Hz. That kills desk rumble and traffic without touching anything you want. Then look for mud around 200 to 400Hz and cut 2 to 3dB if the sound feels boxy.
For harsh sibilance, a narrow cut around 5 to 7kHz of 3 to 4dB usually does it. Use a de-esser only at gentle settings, since aggressive de-essing on whispering strips the exact frequencies the whisper lives in.
Compression: 2:1 ratio, attack around 20ms, release around 200ms, and no more than 3 to 4dB of gain reduction on the loudest parts. The slow attack matters because it lets the initial transient through before the compressor clamps down. That transient is the tingle.
Finish with a limiter set to a -3dB ceiling. It is a seatbelt, not a sound.
Binaural editing, the part nobody explains
If you recorded with a binaural mic like a 3Dio, do nothing to the stereo image. Any mono-summing plugin, stereo widener, or single-channel noise gate will collapse the spatial cue you paid money to record. That collapse is why some expensive setups still sound flat on headphones.
If you recorded mono and want spatial movement, duplicate the track, pan the copies hard, and offset one side by 5 to 15 milliseconds to fake distance. It is not true binaural, but it reads as movement across the head and beats a static center image.
Process both channels together, always. Applying a plugin to the left with slightly different settings than the right makes the whole thing feel unstable and slightly nauseating in headphones.
Editing on your phone
CapCut on a phone handles the entire ASMR workflow: cutting, volume keyframes, basic noise reduction, and export. Plenty of creators edit that way 99% of the time and their videos sound fine. The limitation is fine EQ control, not quality.
The phone-only trick worth knowing is volume keyframing instead of compression. Manually riding the volume up in the quiet passages gets a more natural result than any mobile auto-level. It is the mobile version of clip gain.
If you are recording on a phone with no external mic, record conservatively and boost in editing. Amplifying audio in post gives you room to undo a mistake. A clipped recording is gone forever.
Who this workflow suits, and who should skip half of it
If you are making long sleep videos over 30 minutes, the priorities are consistent volume and zero sudden changes, so spend your time on clip gain and skip the fancy EQ.
If you are making short trigger-focused clips for TikTok, spend the time on transient clarity and export loud enough to survive platform normalization.
This approach works best for people who can listen critically for two hours without their ears fatiguing. If you are editing tired at 1am you will over-process, guaranteed. Come back in the morning and you will hear the damage immediately.
Export settings by platform
YouTube: 48kHz audio, AAC at 384kbps if the encoder allows, video at 1080p or 4K, peaks at -3dB. YouTube normalizes toward -14 LUFS and mostly leaves quiet ASMR alone, so do not crush it chasing a loudness target.
TikTok and Shorts: same sample rate, but expect heavier compression on their end. Leave more headroom and do not rely on very quiet detail surviving. Anything below about -40dB disappears on phone speakers.
Audio-only platforms: export a separate WAV at 48kHz 24-bit rather than ripping audio out of your compressed video file. Ripping a second-generation encode is where a lot of muddy podcast versions come from.
Frequently asked questions
What software do ASMR creators use to edit their videos?
Most working ASMR creators use one of four: DaVinci Resolve on desktop because it is free and its Fairlight audio page handles clip gain and EQ properly, Adobe Premiere with Audition for repair work, Audacity for audio-only editing, or CapCut on a phone. The software matters far less than the settings. A careful edit in free Audacity beats a heavy-handed edit in a paid suite every time, because the real failures in ASMR editing come from over-processing, not from missing features.
How do you remove background noise from ASMR recordings without ruining the sound?
Use a noise profile taken from actual silence in your own recording, never a preset, and keep the reduction amount around 6dB. Anything past 10 to 12dB starts removing the faint high frequency detail that makes ASMR trigger, producing the watery, artificial sound creators complain about. Do your manual clip gain first so the noise reduction has less work to do. If the room is genuinely noisy, record at night with the HVAC off instead. Fixing the room beats fixing the file.
How long does it take to edit an ASMR video?
Around 2 to 5 hours for a finished 30 minute video once you know the workflow, and 6 to 8 hours the first several times. Manual clip gain is the bulk of it, since smoothing every protruding sound by hand is slow work no plugin does well. Creators who batch their editing save real time by recording three or four videos in one session with identical mic placement, then reusing the same EQ and compression chain across all of them.
How do you sync audio and video in ASMR editing?
Clap once on camera at the start of every take, then line up the waveform spike from the external recorder with the spike on the camera audio. Once aligned, group or nest the clips so they cannot drift apart when you cut. Sync drift is the most common thing that ruins a finished long-form ASMR edit, because a 40 minute file recorded on two devices with slightly different clocks can wander by several frames by the end. Check sync at the start, middle, and end before exporting.
Should ASMR videos have background music or added sound effects?
Usually no. Background music competes with the trigger sounds for the same attention and masks the quiet detail people came for. If you want ambience, a very low bed of rain or brown noise sitting 20 to 25dB below the triggers works, and even then some viewers will ask you to remove it. Added sound effects almost always read as fake, because listeners can hear that the reverb tail does not match the room the rest of the video was recorded in.
The Noise Reduction Ceiling and Why It Destroys Tingles
Most DAWs default to aggressive noise reduction settings that were designed for podcast speech, not ASMR. The problem is that ASMR triggers live in the same frequency range as the room noise you're trying to remove.
Gentle tapping, fabric rustling, and whisper consonants all have energy above 4kHz. Standard noise reduction pulls a broad cut across that range, and the result is a muffled, underwater quality that kills the crisp transient detail that produces tingles.
In iZotope RX or similar tools, set your noise reduction reduction amount to 6 dB or below and use a gentle curve rather than hard cutoffs. You want to lift the noise floor slightly, not eliminate it. A little room ambience actually helps convince the listener they're in a real space with another person.
The better long-term fix is acoustic treatment on your recording space. A quiet room means less noise to remove, which means less processing damage. Furniture, rugs, and moving blankets work. You don't need professional foam to get a noticeable improvement.
Compression Settings That Preserve Transient Detail
Compression is useful in ASMR editing but only at very gentle ratios. The goal is evening out level differences between quiet whispering and louder tapping, not squashing the dynamic range flat.
A 2:1 ratio is the upper limit for most ASMR content. Higher ratios kill the dynamic contrast that makes tapping feel satisfying and whispers feel intimate. The attack time matters as much as the ratio. Fast attack (under 5ms) clamps down on transients before the listener hears them, removing the initial punch of a tap or scratch.
Set attack to 20-30ms for tapping content and 10-15ms for whisper-focused content. This lets the leading edge of each sound through before the compressor engages. Release time between 200-400ms works for most ASMR. Faster releases cause pumping artifacts that are audible and distracting.
For multi-trigger videos mixing whispers and tapping, consider using two separate compression passes on different frequency bands rather than one broadband compressor. This gives you independent control over how each trigger type is treated.
EQ Moves That Enhance Specific Triggers
Different ASMR triggers have different frequency signatures, and a single EQ curve applied across an entire video rarely serves all of them equally well.
Whispers benefit from a gentle boost between 2kHz and 5kHz where consonants and breath texture live. Too much boost here creates sibilance, so cut a narrow band around 6-8kHz if you're getting harsh S and T sounds after boosting presence.
Tapping sounds are full in the 800Hz to 2kHz range. If taps sound thin or hollow, a small boost here adds body. If they sound boomy or muddy, cut below 200Hz where room resonance builds up.
Paper and crinkling sounds often need a high-shelf boost above 8kHz to bring out the satisfying texture of each crinkle. This is the range most frequently damaged by noise reduction, so if your crinkling sounds flat after processing, this is usually what to restore.
Mouth sounds are the most sensitive to EQ. Even small boosts in the 3-6kHz range can make lip smacks and tongue sounds uncomfortably loud. Treat mouth sounds conservatively and let the microphone placement do most of the work.
Stereo Widening and Binaural Processing Without Artifacts
True binaural ASMR requires a binaural microphone setup at recording time. You cannot turn a mono or standard stereo recording into binaural in post. Software that claims to simulate binaural processing adds phase artifacts that make headphone listening worse, not better.
What you can do in post is add gentle stereo width to standard stereo recordings. The M-S (mid-side) technique lets you widen the stereo image without affecting the mono center where most triggers are recorded. Most DAWs have a stereo widener plugin or you can do it manually with M-S processing.
For realistic spatial movement, pan automation is more effective than algorithmic reverb. Moving a tapping sound slowly from left to right over 10-15 seconds creates a convincing 3D sensation without adding the room-coloring that reverb introduces.
Avoid stereo widening on mouth sounds and whispering. These should stay centered because off-axis mouth sounds are perceptually strange and break the intimacy of the personal attention format. Apply widening only to ambient layers or environmental sounds underneath the primary trigger.
Export Settings That Preserve Audio Quality
Exporting at the wrong settings after a careful recording and editing session throws away a significant portion of your work. YouTube and most platforms apply their own compression on upload, so your export format matters for how much quality survives that second compression pass.
Export as WAV or FLAC at 48kHz/24-bit before uploading to any platform. Do not export as MP3 and then upload, because the platform will compress the MP3 again and you'll get double lossy compression artifacts that are audible on headphones.
For YouTube specifically, upload at the highest quality your file size allows. YouTube targets a certain bitrate per resolution level for audio, and starting from a higher quality source means the final delivered audio is cleaner.
Normalize your final export to -1 dBTP (true peak). This prevents clipping during platform transcoding while keeping your volume competitive with other videos. ASMR content often runs quieter than general YouTube content, and a properly normalized file ensures listeners don't have to crank their volume to uncomfortable levels.