Walkstamp knowledge base, by topic. Each block opens with a click.
If what you are looking for is not here, use the Talk to us button in the corner of the
tool — a problem, an idea or a compliment all arrive in the same place.
Before anything else: Walkstamp runs entirely in your browser tab. The video is not uploaded anywhere, and that is why almost everything here depends on your machine and your browser — not on our network. That explains half the questions on this page. The other half is in Information security.
It watches a screen recording, keeps only the moments where the screen really changed, transcribes what was said, and puts the two together in a document — PDF, Word, HTML, Markdown, slides, SCORM, CSV or JSON. What was a forty-minute video nobody watches becomes twelve screens with the text of what was being said on each one.
It does not keep the video. Once the document is generated, the video no longer exists inside it — and it never existed on any server at any point. It also does not edit video, does not cut clips for republishing and does not make video subtitles: the end product is always a document.
Because there is nowhere to upload to. The browser reads the file from your disk, extracts the images, runs the speech model and assembles the PDF in the tab's own memory. There is no upload at any step, and you can check that in thirty seconds with F12 open on the Network tab — the total sent stays at zero even with a multi-gigabyte video. The step by step is in Information security.
No and no. Open the tool and use it. There is no install, no extension and no sign-up. There is a single file version you download and open with no internet at all, for when the company network blocks everything — it is on the pricing page.
Current Chrome and Edge, on a computer, is where everything works. Firefox and Safari open the tool and generate a document from a video, but they do not have screen capture with system audio or the control window. On a phone you can read and review, but recording a computer screen is a computer's job.
Whole screen. Sharing a window has two traps: the browser does not deliver computer sound in that mode, and everything that happens outside that window disappears from the evidence — including the error pop-up that appeared in the corner. That is why the picker opens on the whole-screen tab, and the “share audio” box comes already ticked. Forgetting to tick that box is the most common cause of a recording with no sound.
Three seconds between you picking the screen and the recording starting. Without them, the first frame is almost always the browser's own screen picker — the window you closed a moment ago — and you only find that out during review. It can be turned off in the options.
“Mark this step” keeps the screen at that instant, even if it has not changed enough for the automatic detector to keep it on its own. Right below the button there is a text box: what you write there becomes the note for that frame. It always comments on the last mark, never the next one — nobody describes a problem before seeing it. The reason it exists is simple: at review time, half an hour later, seven marked frames with none of them saying why are not worth much.
Pausing stops keeping screens and stops transcribing, but the elapsed time keeps running. If the test waited thirty minutes for something to process, the evidence has to show thirty minutes. What pausing avoids is filling the document with identical frames and side conversation. On resume, a screen is kept right away: after a wait, what changed is exactly what you wanted to document.
While you record, the Walkstamp tab sits behind the screen you are sharing — that is, out of your reach. The control window is a small window that stays on top of everything, with the clock, the audio meters, and the buttons to mark, pause, mute the microphone and stop. The comment box is there too. It exists in Chrome and Edge; where it does not, everything keeps working from the tab.
Because keeping video is the opposite of what the tool does. While you record, it is already extracting the frames and transcribing — by the time you press stop, the material is ready. Keeping the video on disk would mean writing gigabytes only to throw them away.
When you stop, a tail of audio is left over that has not yet filled a transcription window. That is the bar you see. While it runs, steps 3 and 4 stay closed on purpose: the frame list is still changing, and editing a list that shifts under you is how a frame gets lost. If you do not want to wait, the button becomes Do not wait.
MP4, MOV and WebM. Size does not matter — the file is not uploaded, so there is no upload limit to blow through. What limits it is your computer's memory, and very long videos are better opened in pieces.
You can pick a video straight from your Drive: the browser downloads the file to your machine and the work happens there, as with any other file. For YouTube, the route is the transcript — paste the text YouTube itself generates into the transcript box.
Select two or more .json files generated earlier and they become chapters of a
single document, in the order you picked. It is how an afternoon of testing with four cases
becomes one report. Time inside each session stays true; the gap between them is not invented.
From a speech model that runs inside your browser. The first time it is downloaded — from 77 MB to around 400 MB, depending on the quality you pick — and it stays in the browser for the next times. The audio does not leave the machine at any point.
You can, and it is the fastest route: whoever has the .vtt or the
.srt of a Teams, Meet or Zoom meeting does not have to download any model. Drag the
file onto the transcript box, paste the text, or use the Open an existing caption file
button right below it. Loose text with no timestamps works too — only then the speech is not
matched to each screen.
The two channels are captured and transcribed separately, and the document says who said what. In a testing session with a user, that is the difference between “someone said” and “the participant said”.
System names, internal acronyms and proper nouns are what the speech model gets wrong most — and they are exactly the words that matter in a corporate document. The vocabulary is a list of your own terms that the tool applies over the transcript once it is done. It sits right below the text box, which is where the error shows up.
Translation also runs in your browser, wherever it has the translator built in. You can
translate into one language at a time and then generate the PDF, the Word or the Markdown as
usual — or tick several languages at once and get a .zip with the same document in
HTML in each of them.
It compares each screen with the previous one and only keeps it when the change is big enough. A cursor that moved is not a change; a field that turned red is. At review time you see every one it kept and discard the ones that are no use — the discarding is yours, the tool only keeps you from starting with three hundred identical screens.
You can redact a region of the image — a tax ID, a name, an amount. The redaction is applied to the image that goes into the document, it is not a layer on top: in the delivered file, what was underneath no longer exists.
Inside each frame's magnifier: rectangle, arrow and pen to point at what matters; crop to take out the useless part of the screen; duplicate when the same screen proves two things (before and after an action); and compare, which overlays two frames with a blend control — it is the answer to “what changed between these two?”, which the eye does not find on its own when a button moved eight pixels.
It is the most repeated gesture for whoever maintains a document: the screen changed, the procedure did not. Instead of attaching, moving it with the arrows and discarding the old one, there is a button that swaps the image of that step and keeps everything else.
A long document can be split into named tasks, and each task becomes a chapter with its steps underneath. You can also generate a separate document per chapter, from a single capture.
PDF is what almost everyone wants: it closes, prints and attaches. Word is for whoever will keep writing on top of it. HTML is a single file, images inside, that opens in any browser. Markdown is to paste into a wiki. Slides to present. SCORM to load into a training LMS. CSV and JSON are data: the CSV goes to a spreadsheet, and the JSON is what brings the document back here later.
Test evidence, a work instruction, meeting minutes and a usability session ask for different fields and speak to different readers. The scenario you pick changes the labels of the identification fields, which fields appear, and the objectives offered in the prompt. “Context for AI” asks for no identification at all: the reader is a machine.
The layout decides how the screens are distributed — automatic, one per page with the speech from that stretch, or a compact grid of six per page. The paper size is A4 or Letter. Both sit in the Format section, before the buttons that produce the file. On a paid account, both can arrive set from the account, and the screen says where they came from — changing them there applies to that document only.
They are the fields that turn a pile of screens into evidence: which test case it is, on which system, which ticket, and whether it passed or not. They are free. What is paid for is having them arrive filled in from your account settings.
The fingerprint (SHA-256) of each image is printed next to it, in the evidence scenario. It serves to prove later that the image was not swapped. Automatic numbering gives the case a sequential number.
Attaching the PDF in a conversation with a model is not enough: it has to know how the document is organised, otherwise it ignores the clock times and mixes what it saw with what was said. The prompt text is generated on its own, tuned to your video and to the objective you picked. It starts collapsed — the normal gesture is to copy and paste, not to read.
That is what the .json is for. Keep it next to the PDF: it brings the whole
document back into the tool — frames, transcript, notes, redactions — and you fix what you need
and generate again. Without it, a typo in a note costs you a re-recording.
The .zip carries the document in several formats at once, with the images loose
— it is the package to archive, or to hand to whoever will work on the images outside here. It
also comes back into the tool.
There is a page that checks whether a generated document matches the fingerprints it declares itself — verify a document. It is for whoever receives the evidence and wants to know whether it has been tampered with.
The tool in the browser is free and stays that way. What the plan pays for is the account: your brand on the document, the defaults that arrive filled in (company, environment, label, layout, paper size), the team's document templates, the test run and the vocabulary. The full list and the prices are on pricing.
The key is checked inside your browser. There is no call to any server: it works with no internet, and we do not find out when or where it is used. The normal route is to sign in through the link that arrives by e-mail; pasting the key by hand is still there for whoever received it some other way.
What the company sets on the account arrives filled in inside the tool — and stops applying the moment the person holding the case changes something. Whoever knows that this document needs a different layout is the one looking at it. A default is a default, not a straitjacket.
On a team account you can upload a spreadsheet of test cases and hand out the links: each person gets an address that opens the tool already filled in with their case, and marks it done when they finish. The control screen shows what was done, by whom and when — with the receipt of the generated document.
The periods are written in the privacy policy, and there is code behind them: a daily routine deletes what is past its period, file by file. Invoices are kept for the statutory tax retention period — that is stated there, with the reason.
It is almost always one of two things. The first: you shared a window — the browser does not deliver computer sound in that mode, only for a tab or the whole screen. The second: the “share audio” box in the picker was left unticked. The tool tells you which of the two it was, and the end screen says whether any channel was silent the whole time.
What happened is that the screen did not change enough — reading a document, for example. In that case, mark the steps by hand while recording. The tool also keeps the reasons for every discarded frame and shows them in the technical report, which goes along when you talk to us.
It comes from a public CDN address, and some corporate networks block it. Two routes: use a ready-made transcript (item 4), or ask IT to allow the domain. If you have downloaded it before and want the space back, there is a delete the downloaded speech model button, next to the speech model selector.
Everything runs in the tab's memory, so a very long video with many screens weighs on it. What
helps: lower the image quality, generate the document in separate chapters, or split the
recording in two and join the .json files afterwards.
The Talk to us button, in the corner of the tool, sends a problem, an idea or a compliment. A technical report from your browser goes with it — versions, available features, what the recording logged — and you see the report before sending it. No video, image or transcript goes in it.