Rate each row against a defined scale.
Turn written feedback into a score with a clear meaning. Define what each level requires so you can sort rows by the evidence they contain.
Try with example dataStart with the three fictional rows below. Review the rules before processing.
100 successful rows free, every job. No account needed. More rows use prepaid credits; no subscription.
Your full spreadsheet stays in this browser. Selected columns and your decision definition are sent to column.help and TypeSafe AI for processing. Data handling ↗
From bug reports to reproducibility scores
The Bug report reproducibility tool runs from 1 (symptom only) to 5 (repeat confirmed). A higher score means better reproduction detail, not a more severe bug.
| Before · your source column | After · your added column |
|---|---|
| Message | Score |
| The export is broken. | 1 · Symptom only |
| The export fails after I click Download. | 2 · Action identified |
| In Chrome 130 on macOS, open Reports with the attached two-row CSV loaded, choose CSV, and click Download. An empty file downloads instead of the two rows. This happens on all three attempts. | 5 · Repeat confirmed |
Define the rules for your rows
- Choose one dimension. A detailed bug report can describe a minor issue; detail and urgency need different scales.
- Describe the evidence required at each level. Use observable requirements, such as steps, inputs, and expected behavior, instead of vague labels like good or bad.
- For cumulative scales, choose the highest level fully supported by the row. Missing details should not earn a higher score. Review uncertain scores before exporting.
The starting definition for this example
- 1 · Symptom only
- No reproduction procedure is stated; at most a symptom is described.
- 2 · Action identified
- An action leading toward the issue is stated, but steps or the observed result are missing.
- 3 · Partial procedure
- Steps and the observed result are stated, but required inputs or prerequisites are missing.
- 4 · Reproducible procedure
- Ordered steps, required inputs or prerequisites, and observed versus expected behavior are stated.
- 5 · Repeat confirmed
- A reproducible procedure, relevant environment or version, and confirmation that repetition causes the same result are stated.
Define what a higher score actually means
Start by naming one property you want to measure from the text. Bug-report reproducibility is a useful example because each level can describe evidence a reader can identify: a symptom, an action, a partial procedure, a complete procedure, and confirmation that the issue repeats. The score describes the report's detail, not the impact of the bug.
Write descriptions for the middle levels as carefully as the endpoints. A scale with 1 labeled Poor and 5 labeled Excellent leaves most decisions undefined. Explain what has to be present for a row to move from 2 to 3 or from 3 to 4. Avoid criteria that require information your selected columns do not contain.
For a cumulative rubric, later levels build on earlier requirements. In the example above, mentioning a browser version does not earn a 5 if the message provides no reproduction procedure. The row must satisfy the lower requirements as well as the additional requirement for the highest level. This makes missing evidence easier to handle consistently.
Use scores to compare evidence across rows
For bug reports, a reproducibility score can help you find reports that need a follow-up question before investigation. A low score indicates missing reproduction detail, while a high score indicates more of the defined evidence is present. Neither score tells you whether the reported bug is real or how quickly it should be fixed.
For feature requests, a rationale score can separate an unsupported preference from a request that explains a workflow and provides a concrete example. The result can support a research review queue. It does not measure customer value, market demand, or the cost of implementing the request unless those are separately defined and evidenced.
For operational requests, specificity scores can reveal what information is missing. A request for a resolution that names an action and a clear acceptance condition is easier to act on than a general complaint. Use the score to decide what to inspect or ask next, and retain the original text for that follow-up.
Check the boundaries between neighboring levels
A useful preview includes examples near the boundaries, not just an obvious 1 and an obvious 5. Compare rows assigned to adjacent levels and ask which requirement separates them. If the distinction is hard to explain, revise the rubric. Adding more levels does not help when the existing levels are already difficult to distinguish.
Do not combine unrelated dimensions into one scale. A message can be polite but vague, urgent but detailed, or strongly worded without describing a serious problem. If you need to evaluate both detail and urgency, use two definitions and keep their results in separate columns after export.
Keep the score separate from confidence. A result of 5 means the row was assigned to level 5 of your scale. It does not mean the tool is certain. Review low-confidence results and spot-check other levels before relying on the distribution. Counts by level are often easier to interpret than an average when the levels do not represent equal intervals.
Should you Categorize, Score, or Flag?
The same message can support different questions. Choose the output you need before choosing a tool.
- Categorize
- Which kind is it? Choose one label from a set, such as Billing or Shipping.
- Score
- To what degree does it meet the rubric? Assign a defined level, such as 1 through 5 for reproduction detail.
- Flag
- Does it meet this condition? Return Yes or No, with uncertain results marked for review.
A message about a broken export might be categorized as Technical, scored for the detail of its reproduction steps, and flagged for an explicit refund request. Keep those questions separate so each result remains interpretable.
Try Score on your spreadsheet
- Add and preview your data. Open a tool below, then paste from Excel or Google Sheets or upload a CSV. Check the header row and column names.
- Choose the source columns. Select the fields that contain evidence for your question. Leave out unrelated information and any existing answer column.
- Define the result. Edit the starting definition and name the new column. Make sure the rules cover the examples and edge cases you expect.
- Preview, review, and export. Preview a sample before processing the remaining rows. Inspect uncertain results, correct mistakes, then copy or download the results for your spreadsheet.
Common questions about Score
Does a score represent a probability?
No. It is a value from the rubric you define. A score of 4 on a five-level scale is not an 80% probability. Confidence is a separate review signal, and neither value is a guarantee.
Can I change the scale?
Yes. Starting tools provide editable levels and descriptions. In Advanced / Edit definition, you can define between two and ten levels with increasing integer values. Keep the distinctions meaningful for the text you are evaluating.
Can I use a score to rank rows?
You can sort the exported score column in your spreadsheet. Rows at the same level are tied under that rubric; their order does not imply a further distinction. Read the original evidence when you need to decide between tied rows.
What if a row does not provide enough evidence?
Define how missing evidence affects the scale. In a cumulative rubric, absent information does not satisfy a requirement for a higher level. The lowest level should describe the absence of relevant evidence when that is appropriate. Completely empty selected inputs are skipped.
Every type follows the same steps: select your source columns, define the result, preview a sample, then review and export. Your full spreadsheet stays in your browser; only selected columns are sent for processing. Data handling