# AI Log 03 — Evaluation Layout and Reflection Structure

## Purpose of This Log

Record AI assistance related to structuring the evaluation section, Alpha usability testing summaries, Before / After refinement cases, testing tasks, result presentation, and checking claims against real project evidence.

## 1. Date / Project Stage

- Date: Week 8-9
- Project stage: Evaluation refinement / portfolio evidence organization
- Related portfolio/system section: Evaluation & Reflection; Testing Overview; Alpha Usability Testing Summary; Testing Tasks; Testing Results; Iterative Refinement; Final Reflection; AI Disclosure

## 2. AI Tool Used

- Tool name: Google Gemini and OpenAI ChatGPT
- Model: Gemini 3.1 pro; ChatGPT 5.5 Thinking
- Access date: Apr. 29, 2026
- Link: https://gemini.google.com/ and https://chatgpt.com/

## 3. Intended Task

The AI interaction was used to support the organization of HoldLight’s evaluation and reflection page.

The goal was to turn testing material into a clear portfolio structure without overstating the strength of the evidence. The page needed to show how the team moved from alpha usability testing to broader evaluation evidence, how findings informed design changes, and how AI use was disclosed responsibly.

Gemini was used to support evaluation-section organization, summary framing, and clearer reflection wording. ChatGPT was used to refine section order and check whether the claims stayed connected to the actual project evidence. AI did not create participant data, decide the findings, or replace the team’s interpretation of testing.

## 4. Prompt

Summarized organization prompt:

"Help me organize the Evaluation & Reflection page for the HoldLight process portfolio.

Context:
HoldLight is a mobile-first accessible climbing support system for blind and low-vision climbers. The evaluation page should combine alpha usability testing, staged testing evidence, testing tasks, results, before/after refinements, final reflection, and AI disclosure.

Requirements:
- Focus on section order, evidence structure, before/after refinement explanation, and cautious wording.
- Keep the content evidence-based and avoid overclaiming final effectiveness.
- Include alpha usability testing with three participants only if this matches the actual project evidence.
- Show testing tasks for route understanding, audio guidance, uncertainty recovery, and comparison with human support.
- Present results in a readable table and summary cards.
- Connect each before/after refinement to a user-observed issue.
- Make the AI reflection transparent: prompt, outcome, team action, verification, and ethics.
- Keep the page mobile-friendly and readable.
- Do not invent testing results, participant quotes, metrics, or claims.
- Clearly label limitations.

Expected output:
Suggest a page structure, section order, card/table layout, and concise copy framework for the evaluation and reflection page. The output should be reviewed manually against the project evidence."

## 5. AI Outcome

Gemini and ChatGPT suggested a staged evaluation structure with sections for testing overview, participant / phase summary, task cards, results, before/after refinements, final reflection, AI-assisted development reflection, and references.

### What worked

- The AI helped turn a large amount of testing content into a clearer page sequence.
- The suggested structure made it easier to separate alpha testing, broader study phases, tasks, results, and reflection.
- The card and table layout ideas helped make evaluation evidence easier to scan.
- The AI suggested linking refinements back to observed problems, which matched the portfolio’s process-evidence goal.
- The AI disclosure structure helped explain prompt, outcome, action, verification, and ethics in a consistent way.
- Gemini was useful for grouping evaluation evidence into a readable narrative sequence.
- ChatGPT was useful for checking whether the AI reflection followed prompt, outcome, action, verification, and ethics.

### What did not work

- Some AI wording sounded too confident and risked implying final clinical or safety effectiveness.
- The first structure did not clearly separate alpha usability testing from broader evaluation evidence.
- Some generated copy was generic and not grounded in HoldLight’s actual testing tasks.
- The AI occasionally suggested metrics or findings that needed to be checked against the team’s real content.
- Some generated summaries sounded too complete and needed manual limitation wording.
- Some suggested evaluation labels had to be checked against the real participant and testing-stage evidence.
- The output needed manual editing so participant numbers, task names, limitations, and safety boundaries stayed accurate.

## 6. Team Action After AI Output

- Modified: We rewrote the evaluation introduction to explain the staged process and the purpose of each phase.
- Modified: We clearly labelled the coursework alpha test and avoided presenting it as final proof of effectiveness.
- Refined: We organized testing evidence into overview, phases, tasks, results, before/after refinements, and final reflection.
- Refined: We connected each refinement to a specific observed issue: overloaded entry flow, verbose audio cues, and slow hold correction.
- Refined: We added explicit boundary and safety notes to avoid overstating what HoldLight can do.
- Discarded: We removed generic success claims and any unsupported AI-suggested findings.
- Kept: We kept the staged page structure and the AI reflection pattern because they made the evidence trail easier to follow.

## 7. Verification Against User Requirements

- Requirement checked: The evaluation page should show whether HoldLight supports route understanding, low-overload guidance, uncertainty recovery, and human fallback.
- Requirement checked: The portfolio should make the design process and testing evidence traceable.
- Requirement checked: AI use should be disclosed without replacing original research or team judgement.
- Manual test / walkthrough: We reviewed the page order from testing overview to final reflection to check whether the evidence story was understandable.
- Manual test / walkthrough: We checked tables and cards for readable labels, section headings, and consistent evidence links.
- Accessibility check: We checked that long testing tables used mobile-friendly labels and that content was not dependent on visual decoration alone.
- Accessibility check: We reviewed whether limitations and safety notes were visible near the results, not hidden at the end.
- Evidence used: `evaluation-reflection.html`; `content-summary.md`; testing images in `assets/images/testing-evidence/`; before/after images in `assets/images/beforeafter/`; AI disclosure links in `ai-logs/`.

## 8. Ethical / Accessibility / Reliability Considerations

- Possible risk: AI-generated evaluation writing may overclaim results or make the prototype sound more validated than it is.
- Possible risk: Evaluation summaries may blur the line between alpha testing, pilot refinement, and controlled initial evaluation.
- Possible risk: Accessibility claims may become too general if they are not tied to BLV user requirements and testing evidence.
- Possible risk: Safety limitations may be underemphasized if the layout focuses only on positive outcomes.

- How the team reduced the risk: We checked participant numbers, tasks, metrics, and limitations against the project content.
- How the team reduced the risk: We labelled alpha testing as an early usability check and kept broader evaluation claims separate.
- How the team reduced the risk: We added boundary and safety notes directly in the results section.
- How the team reduced the risk: We made AI use transparent through the AI reflection section and `/ai-logs` links.

## 9. Final Use in Project

- Used in files/pages: `evaluation-reflection.html`; `testing-iteration.html`; `final-reflection.html`; `portfolio.css`; `portfolio.js`; `content-summary.md`; `assets/images/testing-evidence/testing-contact-chat.jpg`; `assets/images/testing-evidence/testing-interview-chat.png`; `assets/images/testing-evidence/testing-internal-prototype.png`; `assets/images/testing-evidence/testing-participant-climb.jpg`; `assets/images/testing-evidence/testing-final-onsite-setup.jpg`; `assets/images/beforeafter/interfacebefore.png`; `assets/images/beforeafter/interfaceafter.png`; `assets/images/beforeafter/cuebefore.png`; `assets/images/beforeafter/cueafter.jpg`; `assets/images/beforeafter/deletebefore.png`; `assets/images/beforeafter/deleteafter.jpg`
- Final status: Used after manual revision and evidence checking.
- Notes: This AI interaction supported page structure, copy organization, and reflection framing only. The testing evidence, interpretation, and final claims were checked and edited by the team.
