8 comments

  • syntaxing 3 days ago
    I like this so much. Maybe I’m romanticizing but hoping tools like this give local news a chance against monopolies like Sinclair Broadcasting.
    • cyanydeez 3 days ago
      It'll merely generate hyper local bots that generate hyper local slop. The media landscape is downhill.
      • ChrisMarshallNY 22 hours ago
        I think it's just lead generation. If the paper decides to use AI to write/edit stories, that's another topic.
      • binarymax 1 day ago
        Are search summaries slop?
        • ginko 1 day ago
          Yeah
          • binarymax 1 day ago
            It’s not a blanket ‘Yeah’. It’s a ‘sometimes’ and that sometimes depends on use. Distilling down a whole days worth of research is one of the best uses of AI.
            • irishcoffee 23 hours ago
              My 12 year old told me that their friend group says "Oh I AI'd that" when they make a mistake, as a joke. Yet you trust AI to distill down a days worth of information in a 100% accurate way? When pre-teens already understand how unreliably they are?
              • binarymax 22 hours ago
                If I'm getting citations, and the distillation has strong alignment with citations, then it's fine. In other words: the if the summary is a highlight and not a rewrite then I'm happy. I've studied this extensively and have even created metrics around this problem: https://maxirwin.com/articles/llm-rag/
                • irishcoffee 21 hours ago
                  So, if you review the entire dataset? At that point just write the summary yourself.

                  I use "AI" a lot, it is a fantastic tool. We need to stop pretending it is some kind of panacea. It's a tool.

                  • Forgeties79 20 hours ago
                    > I use "AI" a lot, it is a fantastic tool. We need to stop pretending it is some kind of panacea. It's a tool.

                    I have been repeating some variation of this for months. If people would stop acting like it’s THE tech solution to ALL things ALL the time I bet a lot of critics would quiet down. The overhyping has become exhausting. It’s been going on for years.

                    GPT messed up the math for me the other day when I was simply adding 10 durations for a TRT. Couple of HH:MM:SS inputs, annoying to add up and I had it open.

                    It got it wrong, I told it it was wrong, it got it wrong again. Then out of curiosity I provided it the answer, asked for it to confirm it against the original numbers, and it went “you’re absolute right, it’s [original/wrong answer from earlier].”

                    This stuff happens probably 10-15% of the time for me regardless of the model. Not just math, just super simple crap. It’s wild to see at this point after 3-4 solid years of “hyperscaling” and overhyping. And it’s the kind of thing that keeps people like me from buying in beyond the foot or two we’ve stuck in the water.

              • Zambyte 22 hours ago
                What does "100% accurate" distillation even mean? That sounds contradictory. "Good enough" is fundamentally what makes distillation valuable, not perfection.
              • ToucanLoucan 23 hours ago
                I will say, while I agree with the broad strokes of what you're saying, in my experience when you provide an LLM with data to summarize, instead of it needing to go find it, the hallucination rate goes down staggeringly. I've never had a huge hallucination when the material I'm asking to be summarized is provided to it at the time of the request.

                Any further distance than that though, even summarizing the conversation it's aware of up to that point, is dicier.

        • phoghed 22 hours ago
          Too many people use "slop" as a synonym for "AI Generated", so it's becoming useless as a term
          • karahime 12 hours ago
            Because it has nothing to do with describing the quality of the thing and everything to do with trying to induce a flinch reaction.
          • Forgeties79 20 hours ago
            If the work wasn’t so frequently sloppy we’d stop calling it slop.
            • phoghed 20 hours ago
              Turns out slop is incredibly useful, valuable, and tasty
              • Forgeties79 19 hours ago
                Lots of caveats and qualifiers needed here
      • jasonlotito 22 hours ago
        That's a long way of saying you didn't read the article.
  • 1-6 19 hours ago
    My local news station started focusing on issues far across the country to monetize on sentational headlines and clicks... I'm glad there's concern about improving local news.
  • ChrisMarshallNY 22 hours ago
    Pretty cool.

    I guess the downside might be, that someone could be replaced by this, but it seems that this is a perfectly valid application of AI.

    In fact, lead generation seems to be a natural fit for LLMs.

  • mmooss 3 days ago
    > Early on, the team tried to balance quality with the high cost of searching many small, scattered sources. They discovered that prompting can implicitly control search depth and behavior, but only after trial and error.

    I don't understand the cost here. The number of sources should be relatively tiny - it's not a general Internet search. The amount of data scraped would seem to be relatively tiny - how much data is there on suburban-county, PA?

    • linkjuice4all 14 hours ago
      The Philadelphia region accounts for roughly 2% of the U.S. population and often includes southern portions of New Jersey, northern areas of Delaware, and skirts portions of Maryland. Additionally the third circuit court and a variety of state and federal offices are located there so there are many unconnected systems with different owners.

      We as nerds have failed to create a common standard and we have voters have failed to expand FOIA to surface all of the many meetings, documents, agreements, etc that govern our life so we're left with AI scraping tools that likely do a poor job to fill the gap.

    • mrweasel 1 day ago
      For the individual county it's probably not a ton of data, but multiply it up, it's a lot for all the counties in e.g. Philadelphia.

      Normally there is data, meetings, events and so that you wouldn't report on in a newspaper, because it's really only relevant to maybe a few thousand people, but for those thousands it might be really important.

      My city publishes a lot of stuff on their website, it's hard to find, relevant to maybe a thousand people, maybe less. The school board meetings, public utility companies, companies in general all publish massive amounts of information that's never surfaced, but is relevant to those living in the vicinity. Being able to collect all of this, sort it, assess it's relevans and produce hyper local news could be a massive boost for local grassroots movements and participation in local affairs and elections.

      • thephyber 1 day ago
        Agree that the permutation becomes a problem, but the organization should be super simple: the state mandates the use of some standard index file format for all municipalities governments, not unlike how sitemap.xml works for a website index.

        There is a community-run project that standardizes each election's results across states, counties, and precincts. It could be a template for local info releases.

        • nemomarx 23 hours ago
          do you have documentation on that standard index file? I'm not sure I've seen a Pennsylvania gov page explaining it
    • tclancy 1 day ago
      I think it’s more about how many different types of content you’re looking at. Public google calendars, Facebook event pages, personal sites, etc.
  • FL410 1 day ago
    >The result: Scrape evolved from a small experiment to “load-bearing and critical infrastructure”

    Please be satire

    • scottyeager 21 hours ago
      Definitely raised an eyebrow at this one :D
  • dec0dedab0de 1 day ago
    I'm still annoyed at them for shutting down philly.com
  • calvinmorrison 1 day ago